EEVblog® Electronics Community Forum

Electronics => Microcontrollers => Topic started by: peter-h on August 02, 2021, 01:31:11 pm

Title: GCC compiler optimisation
Post by: peter-h on August 02, 2021, 01:31:11 pm
Cube IDE, 32F417.

I've just had a funny one. I have a boot loader which transfers control to base+32k (0x08008000) and that is where the linker script places main.o. And I put main() right at the start.

However, if you used the SWV ITM / SWD debug interface, the compiler was putting ITM_SendChar() at the start of main.o! This is actually a macro doing some inline code.

It looks like the compiler is placing inline code before normal functions. This is news to me; I thought that compilers didn't change the order of functions in a .c file :) Why should they?

I fixed it initially by changing the function to a normal one

Code: [Select]
static void ITM_SendChar_2 (uint32_t ch)
{
  if (((ITM->TCR & ITM_TCR_ITMENA_Msk) != 0UL) &&      /* ITM enabled */
      ((ITM->TER & 1UL               ) != 0UL)   )     /* ITM Port #0 enabled */
  {
    while (ITM->PORT[0U].u32 == 0UL)
    {
      __NOP();
    }
    ITM->PORT[0U].u8 = (uint8_t)ch;
  }
  return;
}

but the proper fix was do create main_stub.c which contains just main() which then calls real_main(), and in the linkfile you put main_stub.o first

Code: [Select]

  /* The rest of the code goes here, loaded at base+32k, starting with a stub and then the real main() */

  .main_stub.o :
  {
    . = ALIGN(4);
    KEEP(*(.main_stub.o))
    *main_stub.o (.text .text* .rodata .rodata*)
    . = ALIGN(4);
  } >FLASH_APP
   
  .main.o :
  {
    . = ALIGN(4);
    KEEP(*(.main.o))
    *main.o (.text .text* .rodata .rodata*)
    . = ALIGN(4);
  } >FLASH_APP
 
/* This collects all other stuff, which gets loaded into FLASH after main.o above */
 
  .text :
  {
    . = ALIGN(4);
    *(.text)           /* .text sections (code) */
    *(.text*)          /* .text* sections (code) */
    *(.rodata)         /* .rodata sections (constants, strings, etc.) */
    *(.rodata*)        /* .rodata* sections (constants, strings, etc.) */
    *(.glue_7)         /* glue arm to thumb code */
    *(.glue_7t)        /* glue thumb to arm code */
*(.eh_frame)

    KEEP (*(.init))
    KEEP (*(.fini))

    . = ALIGN(4);
    _etext = .;        /* define a global symbol at end of code */
} >FLASH_APP
Title: Re: GCC compiler optimisation
Post by: ComradeXavier on August 02, 2021, 02:00:22 pm
It looks like the compiler is placing inline code before normal functions. This is news to me; I thought that compilers didn't change the order of functions in a .c file :) Why should they?
I don't think there's generally any guarantee that object code will be in the same order as the source.

But assuming that that source order generally holds, remember that the compiler's view of a .c file contains everything that was included (the translation unit). So if the compiler emits a body for an inline function in a header included before your .c file's first function, the inline function's object code will appear before your .c file's first function's object code.
Title: Re: GCC compiler optimisation
Post by: gf on August 02, 2021, 02:15:41 pm
No, there is neither a guarantee for the order of functions, nor for the order of global variables.
If you want to place main() at 0x08008000, then you need to put main into a separate section, and place this section at 0x08008000 via the linker script.
Title: Re: GCC compiler optimisation
Post by: magic on August 02, 2021, 02:21:11 pm
A macro doesn't compile to a separate function. Most likely, your ITM_SendChar is a "static inline" function declared somewhere in some header and gets inserted near the beginning of your C file as described two posts above. A separate copy is also similarly inserted into each other C file which includes that header.
Title: Re: GCC compiler optimisation
Post by: abyrvalg on August 02, 2021, 02:41:19 pm
And we are back at square #1. You don't need to use tricks to solve this, spend some time on a better architecture and it will save you much more time and efforts (especially if you plan to revisit this project in 10 years, as you say). There are better approaches suggested by several people in your original thread, just pick one and ask to elaborate.

To your specific question: the only reliable way to place something at known address with GCC is to specify the placement in linker script (either by adding section attribute to main() or by the name of file containing main() - main.o(.text*)). But even if you solve this main() placement the next question will be how to place some ISR at fixed offset because your "main" part doesn't have it's own vector table and all interrupts land in your bootloader. Then you'll add another ISR and run that circle again. But a more simple solution would be to have a separate vector table for "main" placed correctly and don't bother with functions placement at all.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 02, 2021, 04:07:06 pm
This is interesting - thank you.

Yes I have come across .h files generating code, previously.

This one was correctly (I believe) solved by the linker file method. With old simple tools, this was much more obvious. I recall doing a product into which the customer could load an application, linked to run at base of an EEPROM, 16k up, and that also had a similar stub, containing just one jump instruction, which was placed at the top of the module list in the linkfile.

I am not using interrupts in any of this code. Nothing before main() starts enables interrupts.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 02, 2021, 04:20:53 pm
Cube IDE, 32F417.

OK, but it has nothing to do with this actually.

I've just had a funny one. I have a boot loader which transfers control to base+32k (0x08008000) and that is where the linker script places main.o. And I put main() right at the start.
However, if you used the SWV ITM / SWD debug interface, the compiler was putting ITM_SendChar() at the start of main.o! This is actually a macro doing some inline code.
It looks like the compiler is placing inline code before normal functions. This is news to me; I thought that compilers didn't change the order of functions in a .c file :) Why should they?

I don't know where you got the idea that a C compiler would guarantee that the order of "functions" in object code would have anything to do with the order you put them in source code. Nothing of this kind has ever been guaranteed.

As said above, the only way to control code location in the final object code is to specify this at the link step, customizing the linker script.
You can control where code would go putting it in a dedicated section.

For instance, the typical way of defining the location of the "startup code" would look like this: (first section in the linker script)
Code: [Select]
.text.startupcode :
{
. = ALIGN(4);
KEEP(*(.text.startupcode))
. = ALIGN(4);
} >INSTRMEM
Title: Re: GCC compiler optimisation
Post by: ataradov on August 02, 2021, 05:29:17 pm
Your architecture is not correct. Just make the bootloader read the reset vector address from the binary vector table and you won't have to rely on fixed addresses and will not have to force the compiler to do unnatural things. This is how all bootloaders for ARM work.

Also, you can't just jump to a random address, you need to read out initial SP value anyway. Otherwise you have a potential to run into all sorts of issues as you recompile things.

Here is a typical code that bootloaders use to run the application:
Code: [Select]
static void run_application(void)
{
  uint32_t msp = *(uint32_t *)(APPLICATION_START);
  uint32_t reset_vector = *(uint32_t *)(APPLICATION_START + 4);

  __set_MSP(msp);

  asm("bx %0"::"r" (reset_vector));
}
APPLICATION_START is the address of the application image in the flash (address of the vector table).
Title: Re: GCC compiler optimisation
Post by: peter-h on August 02, 2021, 05:33:42 pm
In some ARM compilers, it appears from looking around the web, you can specify a particular function to be placed at a particular address. But apparently not, as far as I can find, with GCC. A while ago I spent ages on it and nothing worked.

If this was possible it would be a neater way around it, than splitting off the module entry point into a little stub.c file and then putting stub.o at a particular address in the linkfile.

In GCC you can locate a RAM buffer that way, or a general variable, but not a function.

Anyway, that this happens only when ITM_SendChar is invoked, is a gotcha, because it will happen only when debugging.

I don't understand why the stack pointer is relevant. You can set the SP to, say, the top of CCM, and it should not need to be touched through the startup, copying the boot loader to RAM, running the boot loader in RAM, etc.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 02, 2021, 05:38:10 pm
With -ffunction-sections you can place individual functions anywhere you like in the normal linking process. Standard GNU LD does not support flow around, so addresses still must be incrementing monotonically. It is hard to place things in the middle of other tings. GOLD linker is supposed to address this, but I'km not sure of its progress or usability in real life.

SP is relevant because application and bootloader SPs are generally different. It is a bad idea to expect them to be the same. And even if you make them the same, if you are jumping to the application without modifying the SP, then the stack space already used by the bootloader at the time of call will be lost in the application.

You are going against established industry practices without gaining anything in return, but long term troubles. And also short term troubles, given that you have to fight the compiler and force it to do strange things.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 02, 2021, 06:03:31 pm
You can place any function in any section with just an attribute with GCC (and I guess Clang should support it too.) It's really as simple as it gets.

Code: [Select]
__attribute__((section("SomeSection"))) int SomeFunction(int n)
{
        return n*2;
}

No need to create a dedicated object file.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 02, 2021, 06:13:09 pm
But you can't easily place it at a fixed address in the middle of the rest of the code. Not that dedicated object file helps with that either.

The whole issue stems from the wrong architecture. Just doing things the right way solved all of this instantly.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 02, 2021, 06:22:24 pm
But you can't easily place it at a fixed address in the middle of the rest of the code. Not that dedicated object file helps with that either.

Not that it makes any sense as you said, but you can. You just need to lay out sections in the linker script appropriately, and put code in the respective sections as needed.
For just an entry point, it's OK and usually how things are done. Most often, you just have two code sections: a 'startup' section (whatever you name it), and the 'text' section, where the rest of the code goes. Nothing prevents you from having more than two sections. It would just be annoying to maintain. It would just be a very convoluted way of achieving what the OP wanted to achieve.

Title: Re: GCC compiler optimisation
Post by: ataradov on August 02, 2021, 06:27:53 pm
That's why I said "can't easily". With split address space like this you would have to do linker's job manually. Nobody would realistically do that in real projects, it is just stupid.

So fixed things are just places art the beginning or the end of the address space.

Linker in XC32 can do that automatically, you just specify the address for the function though an attribute, but it makes everything else so annoying that it is not worth it overall.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 02, 2021, 07:32:46 pm
" if you are jumping to the application without modifying the SP, then the stack space already used by the bootloader at the time of call will be lost in the application."

I don't understand that. I have the SP at top of CCM so there is 64k stack space.

The boot loader gets copied to base of the 128k RAM.

I don't see a problem with stack operation. In fact it looks like the code which transfers control to the RAM-resident boot loader could even pass function parameters to it, which is done via registers or the stack (if passed by value).
Title: Re: GCC compiler optimisation
Post by: JOEBOBSICLE on August 02, 2021, 07:42:08 pm
Have a look at your stack pointer when you jump (sp register). Is it exactly the same as your original stack pointer when you started into the reset vector? Chances are it won't be.



Title: Re: GCC compiler optimisation
Post by: ataradov on August 02, 2021, 08:13:41 pm
I don't understand that. I have the SP at top of CCM so there is 64k stack space.

Look at how the stack looks like at the time you pass control to the application:
Bootloader reset vector is called, local variables from the reset vector are on the stack .
main() is called, main() local variables and return address to the reset handler  are on the stack
some intermediate functions are called (the ones that decide that it is time to run application), their variables are return addresses are on the stack.

All of those functions have a potential to keep stuff on the stack. Now you jump to the application without resetting the stack pointer and you have lost space at the end.

A bigger problem is that you are committing yourself to having the same SP for the remainder of the product life. You can't decided at a later time that CCM is more valuable for something else and you need to place the stack at some other location. You can switch it in the application, of course, but why even create this headache?

If you want to do things "your way", it is fine, but don't you think that the amount of topics you create with things going wrong is indicative of the fact that you are doing something wrong?

And yes, set a breakpoint at the entry point to your application and see how much stack space you lose exactly though that process.
Title: Re: GCC compiler optimisation
Post by: abyrvalg on August 02, 2021, 08:25:25 pm
I am not using interrupts in any of this code. Nothing before main() starts enables interrupts.
But what will happen after bootloader finishes it’s work, main() starts and enables interrupts? All interrupts will go into your bootloader (where the vector table is). How are you going to direct them to the main part? You’ll either need to do your fixed placement trick again for all ISRs in the main part (so vector table in bootloader could point to them) or provide some kind of entry point table in the main part (reinventing the wheel a vector table).
Title: Re: GCC compiler optimisation
Post by: ataradov on August 02, 2021, 08:39:16 pm
Oh, yes, I assumed that application interrupt table follows that strange startup. With this approach interrupts will not work at all.

You can setup a new interrupt table though SCB->VTOR, but there must be a table somewhere in a first place.

Also, you can't simply jump to the application main(), you need to have startup code that copies initialized data and resets the BSS section. Note that this code is completely separate from what the bootloader is doing.

You need to treat the application as its own entity. Your system must remain operational if the whole bootloader is just a simple stub that jumps to the application. If this is not the case, you are setting yourself up for a failure.
Title: Re: GCC compiler optimisation
Post by: lucazader on August 02, 2021, 08:46:21 pm
Since you are using the CubeIDE you can follow a very similar method to what we do with all of our ST based products that use a custom bootloader:

in the linker script for your application, set the start of flash to be the location the application will be saved to (in your case 0x8008000)
Code: [Select]
/* Specify the memory areas */
MEMORY
{
RAM (xrw)      : ORIGIN = 0x20000000, LENGTH = 128K
FLASH (rx)      : ORIGIN = 0x8008000, LENGTH = 512K - 32K
}

Then in your bootloader you can create a jump to application function wich will do a few things:
De-init any peripherals you were using in the bootloader.
You may not have to do this but it is good practice to as the hal expects peripherals to be in a power on reset state on startup, and other states can cause issues. (but not all the time)
Then it remaps the vector table to your application and sets the main stack pointer.
Then it jumps to the application:

Code: [Select]
void flash_updater_jump_to_application()
{
    platform_specific_hal_deinit(); // de-init all peripherals here

    SysTick->CTRL = 0;
    SysTick->LOAD = 0;
    SysTick->VAL = 0;
    SCB->VTOR = APPLICATION_ADDRESS;
    __set_MSP(*(__IO uint32_t *)APPLICATION_ADDRESS);
    __asm volatile(" mov r1, %0" ::"r"(APPLICATION_ADDRESS));
    __asm volatile(" ldr r1, [r1, #4]");
    __asm volatile(" bx r1");
}

an example of the hal deinit could be something like this (but obviously depends on the peripherals used by your bootloader)
Code: [Select]
void platform_specific_hal_deinit()
{
    HAL_UART_DMAStop(&COMMS_UART);
    HAL_UART_DeInit(&COMMS_UART);
    HAL_UART_DeInit(&TRACE_UART);
    HAL_CRC_DeInit(&hcrc);
    HAL_TIM_Base_Stop_IT(&htim6);
    HAL_TIM_Base_DeInit(&htim6);
}
Title: Re: GCC compiler optimisation
Post by: ataradov on August 02, 2021, 09:01:21 pm
I would not do the deinit stuff. It is much easier to set some flag that application needs to be run on the next reset, then request a device reset. If on a reset you see the flag is set, then just run the application without initializing anything else. This way everything is guaranteed to be reset regardless of things you may forget to reset, or later additions.

I typically use first 4 words in the SRAM. I fill them with some known values to reset the application run. But there are other options for that too.
Title: Re: GCC compiler optimisation
Post by: lucazader on August 02, 2021, 09:08:36 pm
For sure, that works if the bootloader uses the same config for peripherals as your main application.

In our case this isnt what happens and we have had some weird system crashes even after re initing all the require peripherals in the application.
For some of our boards it worked fine to not de-init, but others it was a nightmare to track down the issues.
Title: Re: GCC compiler optimisation
Post by: gf on August 02, 2021, 09:09:15 pm
Also, you can't simply jump to the application main(), you need to have startup code that copies initialized data and resets the BSS section.

Usually it is done in startup code outside main, but you could also do that with a memcpy() and memset() at the very beginning of main().
Title: Re: GCC compiler optimisation
Post by: peter-h on August 02, 2021, 09:17:33 pm
This thread has digressed into boot loaders :)

My boot loader exists only for programming the CPU FLASH (with a field-installable application module, or in hopefully rare cases with a complete replacement for the original factory code; another story). If one is not flashing the CPU, there is no "boot loader". Why have one? The whole thing just starts up at 0x08000000, sets up the hardware, and starts up the RTOS.

If the boot loader is used (for CPU flashing) then after it has done that, it reboots. It can't do anything else because it got loaded into the main RAM and crapped over stuff which would otherwise be running in there. The stack will also have some boot loader related crap on it. Originally I was loading it at RAM base 0x20000000 and it worked fine, except that this address is also where some data+bss for the preceeding code (the boot block) gets loaded (it has to go "somewhere") and since the loader was crapping over that, it was unable to call any boot block functions which relied on initialised data or bss. So I now put the loader halfway up the RAM, where there is absolutely nothing, it has the 64k CCM stack to play with, and it could even use most of the RAM underneath it. It's actually quite simple. The key is that it always reboots (with HAL_NVIC_SystemReset); the RAM resident code never returns anywhere.

And interrupts are not enabled at all in the loader. That complicates things too much because you have to switch the ISRs to RAM as well. .

Unless this CPU is doing something totally weird I can't see why the SP should be affected. It gets initialised as the first instruction in the .s startup.

The loader code does use the same hardware functions as the rest of the product does normally, but in any case I am doing a reset after it has run.

Using the RTC SRAM for storage is a cunning trick and we use it to store a magic number which, if not present, causes the RTC to be initialised to some sensible values (Monday 1st Jan 2021 or whatever). The problem is that it is supercap powered and the supercap might not be charged if you do weird power up/down stuff. So I am using a few bytes in a 4MB serial FLASH chip to store flags which tell the CPU to enter the boot loader module on the next boot-up, and copy stuff in the serial FLASH into the CPU FLASH. I have all this working, with lots of verification, but I am stopping just short of actually flashing the CPU until I have 100% tested it all. I posted some Q here
https://www.eevblog.com/forum/microcontrollers/32f417-best-way-to-program-the-flash-from-ram-based-code/new/#new (https://www.eevblog.com/forum/microcontrollers/32f417-best-way-to-program-the-flash-from-ram-based-code/new/#new)
But perhaps you meant using the normal SRAM for this; I would not risk that not getting corrupted by a reset, and in any case it gets wiped by the startup .s code (zeroing BSS etc).
Title: Re: GCC compiler optimisation
Post by: ataradov on August 02, 2021, 09:19:34 pm
For sure, that works if the bootloader uses the same config for peripherals as your main application.
What dio you mean? The software requested MCU reset in my scenario would reset all the peripherals to their default values.

There will never be any issues, unless reset of the MCU does not fully reset the peripherals, but that would be stupid and I don't think this ever happens for practical devices.

Title: Re: GCC compiler optimisation
Post by: ataradov on August 02, 2021, 09:24:07 pm
And interrupts are not enabled at all in the loader. That complicates things too much because you have to switch the ISRs to RAM as well. .
That is fine. Where is your application vector table located?

Unless this CPU is doing something totally weird I can't see why the SP should be affected. It gets initialised as the first instruction in the .s startup.
Why? You are doing something really weird, but whatever, it is up to you what you do with your code. Why even have assembly startup? Designers of Cortex-Mx cores worked hard to make it possible to program the whole thing in C. Why go back to stone age?

But perhaps you meant using the normal SRAM for this; I would not risk that not getting corrupted by a reset, and in any case it gets wiped by the startup .s code (zeroing BSS etc).
There is no risk, SRAM is guaranteed to retain values over resets as long ad MCU is powered. It is the bootloader code that would be looking at that value, just don't erase it in the code. There is no problem here.

I don't understand. Do you erase the bootloader on the first update? Why even have it in a first place then?
Title: Re: GCC compiler optimisation
Post by: lucazader on August 02, 2021, 09:28:00 pm
What dio you mean? The software requested MCU reset in my scenario would reset all the peripherals to their default values.
There will never be any issues, unless reset of the MCU does not fully reset the peripherals, but that would be stupid and I don't think this ever happens for practical devices.

You mean using something like:
Code: [Select]
NVIC_SystemReset()
I hadn't thought to use that, will it boot to the application at 8008000?
If so that sounds much easier! ill give it a shot, thanks!
Title: Re: GCC compiler optimisation
Post by: ataradov on August 02, 2021, 09:33:15 pm
I hadn't thought to use that, will it boot to the application at 8008000?
It will boot as normal, but you can do the check first thing in the bootloader before anything is intialzied.

To request application run:
Code: [Select]
static uint32_t *ram = (uint32_t *)HMCRAMC0_ADDR;

    ram[0] = 0x...; // Any random values
    ram[1] = 0x...;
    ram[2] = 0x...;
    ram[3] = 0x...;
    NVIC_SystemReset();

And then the first thing in the bootloader main() (or even startup code if you want):

Code: [Select]
  if (0x... == ram[0] && 0x... == ram[1] &&  0x... == ram[2] && 0x... == ram[3]) // same values as before
  {
    // jump to the application
  }

I've used this method many times without any issues.

Furthermore, it is a very convenient way to request a bootloader from the application. Application puts a different set of values and resets the MCU. Bootloader runs, checked the values and understands that it needs to stick around.

There is obviously a chance that SRAM will randomly have the same set of values, but probability is very low, and a power cycle will solve the issue if that ever happens.
Title: Re: GCC compiler optimisation
Post by: gf on August 02, 2021, 09:53:19 pm
Or vice versa, jump to the application by default, unless ram[0...3] indicate that an update needs to be done.

Code: [Select]
  if (0x... == ram[0] && 0x... == ram[1] &&  0x... == ram[2] && 0x... == ram[3] || !app_flash_checksum_is_valid())
  {
    // write new application to flash

    // clear update marker
    ram[0] = ram[1] = ram[2] = ram[3] = 0;
  }

  // jump to the application

Edit: Btw, if the bootloader does not overwrite itself when flashing the new application, then there is also no need to run it from RAM.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 03, 2021, 06:09:59 am
"Btw, if the bootloader does not overwrite itself when flashing the new application, then there is also no need to run it from RAM."

That is an interesting angle. It has been claimed in various places online that the 32F4 does not crash if you program the CPU FLASH with code running out of CPU FLASH. You just get a wait state which persists for the duration of the programming cycle.

I wonder if anyone has actually tested this?

Notably, the ST functions for CPU FLASH programming run out of RAM. Maybe they do it as a general case so the CPU can reprogram all its FLASH, or maybe they know something...

NVIC_SystemReset() should boot to the bottom reset vector, which is in the table at 0x08000000. I never change that table. The discussion mentioning 0x08008000 (base+32k) is merely to do with my application where I retain the bottom 32k across CPU flashing (in most cases) while loading app code at base+32k whose entry point is required to be at 0x08008000. This thread started with an example where this was failing because the compiler was sneaking some other code in before that, pushing the real entry point down a hundred bytes or so.

"Where is your application vector table located?"

In the bottom; never changes. The application runs under the RTOS.

"Why even have assembly startup? Designers of Cortex-Mx cores worked hard to make it possible to program the whole thing in C. Why go back to stone age?"

That is what ST supply, with their development board which we (like most people, I am sure) used as a starting point. I know assembler so am perfectly comfortable with it. It also ensures one doesn't get the very thing which started this thread :) A .s file will never get its modules re-ordered by the assembler.
Title: Re: GCC compiler optimisation
Post by: JOEBOBSICLE on August 03, 2021, 06:27:36 am
You don't need to copy the functions to RAM, you can definitely program other parts of flash using the normal hal functions.

Title: Re: GCC compiler optimisation
Post by: peter-h on August 03, 2021, 06:42:53 am
OK, so it is only for programming the same block (16k or whatever size) that you need to move the programming code elsewhere. This (that someone has actually tested it) would have been useful to know when I was posting this
https://www.eevblog.com/forum/microcontrollers/how-to-create-elf-file-which-contains-the-normal-prog-plus-a-relocatable-block/ (https://www.eevblog.com/forum/microcontrollers/how-to-create-elf-file-which-contains-the-normal-prog-plus-a-relocatable-block/)

I did know that the 2MB version of the 32F4 can program one 1MB bank while executing from the other 1MB bank, but we aren't using that version.

And the RAM resident flash code also covers reprogramming the whole FLASH.

I wonder how ST deal with stuff like timer ticks for timeouts (which imply ISR usage) when getting a few milliseconds' worth of wait states?
Title: Re: GCC compiler optimisation
Post by: ataradov on August 03, 2021, 06:44:05 am
If you execute write operations while running from the same flash bank, code execution will simply stall. This is a well documented behaviour. The reason to run from SRAM is if you want to still have interrupt vectors running (for example to receive next block of data while current one is written). ST code is just generic to address this scenario, there is no conspiracy, they don't "know something".

Your explanation of the memory layout makes no sense. Why is it required to be at 0x08008000? There is no reason for it. You seem to completely misunderstand how Cortex-Mx applications with bootloaders are supposed to be structured.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 03, 2021, 06:47:00 am
" code execution will simply stall."

That's good to know, but it doesn't happen with anything I have used previously. It would just crash, because the FLASH was not readable during a programming cycle, so the CPU fetched a duff opcode and crashed.

"Why is it required to be at 0x08008000? "

It's a complicated explanation, to do with loading an application module "somewhere".

I solved this function reordering business by using a stub containing just one function, and no include files, and linking this in the linkfile at 0x08008000, so there is now no chance of something else getting before that.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 08, 2021, 06:05:52 am
This is an interesting compiler optimisation Q.

This is from the ST libs for flashing the 32F4 CPU FLASH

Code: [Select]
**
  * @brief  Program word (32-bit) at a specified address.
  * @note   This function must be used when the device voltage range is from
  *         2.7V to 3.6V.
  *
  * @note   If an erase and a program operations are requested simultaneously,
  *         the erase operation is performed before the program one.
  *
  * @param  Address specifies the address to be programmed.
  * @param  Data specifies the data to be programmed.
  * @retval None
  * Waits for previous operation to finish
  *
  */

static void B_FLASH_Program_Word(uint32_t Address, uint32_t Data)
{
// wait for any previous op to finish
while(__HAL_FLASH_GET_FLAG(FLASH_FLAG_BSY) != RESET);
// clear program size bits
CLEAR_BIT(FLASH->CR, FLASH_CR_PSIZE);
// reload program size bits
FLASH->CR |= FLASH_PSIZE_WORD;
// enable programming
FLASH->CR |= FLASH_CR_PG;
// write the data in
*(volatile uint32_t*)Address = Data;
}

It is the use of "volatile". They actually use a macro (being ST ;) ) called __IO, which one can understand for writing (not reading, surely) IO pins. But for writing FLASH memory? How can the compiler know that location might not be read back much later?

That func is a stripped down version of the real one, which is full of error condition code which basically cannot happen unless the silicon is defective. In my use I anyway do a verify and then potentially 1 more attempt. The above is working code.

The code is still stupidly convoluted e.g. the use of CLEAR_BIT but that's another story. I wonder about the mentality of the people who write this auto generated code.
Title: Re: GCC compiler optimisation
Post by: abyrvalg on August 08, 2021, 07:33:01 am
volatile is there to force the order of operations I think. Not using volatile would mean the result of the assignment will not be seen by other code until return, so the optimizer could place it in any part of the function.
Title: Re: GCC compiler optimisation
Post by: Siwastaja on August 08, 2021, 07:56:13 am
volatile is there to force the order of operations I think. Not using volatile would mean the result of the assignment will not be seen by other code until return, so the optimizer could place it in any part of the function.

This - for the same reason, all the above control registers are already qualified volatile in the header files.

It's very ineffective to clear and set the PSIZE bits for each word, though. A function writing a whole buffer of words would be much more sensible than calling this function in a loop.
Title: Re: GCC compiler optimisation
Post by: DavidAlfa on August 08, 2021, 08:07:20 am
volatile means that the data can change at any time regardless of the code, so avoid any assumptions when optimizing.
So the compiler doesn't optimize the code based on its value.

For example, this will translate as a while(1) and never work, no matter the value of flag, because the compiler only sees flag=1:
Code: [Select]

uint8_t flag;

void something(){
    flag=1;
    while(flag);
}

void interrupt(){
  flag =0;
}


In this case, as flag is changed by an interrupt, it breaks program workflow, so it must be declared as volatile. Now the compiler will check it in every loop.
Same with anything related to peripherals, ports...

That's one of the first things you learn when programming. When you see how a variable is 0 but he code is still stuck in the loop... you blame everything! Must be this hell of compiler full of bugs!
Happens only once in life and never forget. I still remember that moment when I started fiddling with C 12 years ago. Lost some hair before discovering it :D
Title: Re: GCC compiler optimisation
Post by: Siwastaja on August 08, 2021, 08:26:38 am
We know that, but in this example the reason for volatile is different: it's used in writes only, and here it simply enforces the order of those writes.

(Doing the read-modify-writes with a volatile-qualified control register like in that code is crappy. ST usually does better, if you look at their functions usually they use a temporary variable to avoid multiple unnecessary read-modify-write operations. Personally, I prefer to avoid read-modify-writes of peripheral registers and do only writes, with two advantages: setting the full peripheral state to a known value at once, and with better performance, too. But sometimes this isn't an option as it would create new dependencies between different pieces of code.)
Title: Re: GCC compiler optimisation
Post by: DavidAlfa on August 08, 2021, 10:34:25 am
Yup, sorry if I seemed condescending.
What I meant is volatile is for everything that behaves differently from obvious.
Anyways, these peripherals requiring specific patterns (ex. Flash writing unlock) usually have a macro or function un assembly.
Title: Re: GCC compiler optimisation
Post by: Siwastaja on August 08, 2021, 10:45:03 am
Didn't seem condescending, no probs.

You definitely don't need to write anything in assembly to write the flash, and (static inline) functions are really cleaner than macros.

Just a few simple control register writes like with any peripheral, and then normal memory write operations (qualified volatile to ensure right order of operations), i.e., assignments.

Flash peripherals vary from device to device as usual with ST, but basically the sequence is,
* Write magic numbers to unlock register,
* Possibly set some options like programming parallelism size, allowing higher speed if you have enough supply voltage available
* Write erase command bit to '1' in control register to erase, poll some completion flag
* Write write enable bit to '1' in control register
* Do a normal memory write access, poll for some completion/busy flag.
* Repeat the previous item until finished.
Title: Re: GCC compiler optimisation
Post by: abyrvalg on August 08, 2021, 12:54:46 pm
BTW, regarding the read-modify-write, why nobody uses the Cortex-M’s bit banding (accessing individual SRAM/IO bits via aliased 22xxxxxx/42xxxxxx memory regions)? Seen it only once I guess - one particular PVD bit accessed via hardcoded 42xxxxxx address in F103 SPL. It would be cool to have it supported by compilers (i.e. you declare a reg/var as a struct with bit fields and compiler uses the best access method), perhaps that’s why nobody uses it - no obvious ways to use, no examples. I’ve defined a BB(reg, bit) macro doing the necessary conversion and using it sometimes. One important aspect of bit banding is guarantied atomic write (locked AHB cycle, even DMA couldn’t steal the bus between read and write).
Title: Re: GCC compiler optimisation
Post by: Siwastaja on August 08, 2021, 01:25:27 pm
Maybe because that would be most useful with GPIO, and vendors already offer atomic set/clear registers such as BSRR on STM32, so bit banding is just a duplicated facility for the same.

In peripherals, it's quite rare that you actually need RMW for the control registers; you are in the command of the peripheral so you can just write the register fully with the state you want, no read involved. Sometimes RMW is just for convenience to preserve modularity in your program, think about enabling clocks on RCC. If you want to enable SPI5, you don't want to accidentally turn SPI4 off, so you do RMW just turning the thing you need on. In such cases, performance and code size penalty is completely negligible.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 09, 2021, 06:59:53 am
Would something like this get optimised (a read of CPU FLASH)

       uint8_t* p = (uint8_t*) 0x08000000;
      ch=*p++;

on the assumption that the compiler has not seen any code writing to that location? That would be bizzare, surely?
Title: Re: GCC compiler optimisation
Post by: ataradov on August 09, 2021, 07:17:14 am
It depends on the compiler. Most existing compilers will not optimize direct address casts like this. But you are setting yourself up for failure in the future when a new version of a compiler does this optimization. There is nothing really stopping them from doing it, apart from non-zero amount of poorly written code that will break. Don't be a part of the problem.

Also, there is another possible optimization here - if the compiler knows what would be placed at that address, it may use a fixed value known at compile/link time. And it could just remove the whole section of the code. This would break anything that was added after the compilation (like image size or CRC). Again, this is the case of things changing outside of the compiler control.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 09, 2021, 08:25:36 am
Interesting...

It does make me wonder whether these optimisations have any impact whatsoever on system performance. I have decades' experience of assembler, and all the tricks people used to do (including self modifying code, which I avoided), so I understand this stuff at the machine level. And in most systems some 1% of the code is speed critical, and one generally gains far more there by sitting down and thinking about doing that job differently, than by rewriting it in the slickest assembler possible.
Title: Re: GCC compiler optimisation
Post by: DavidAlfa on August 09, 2021, 09:05:05 am
For sure you use a lot of rmw in peripherals.
Just to enable it por example.
You usually set up the registers, then you set the enable bit.
So you need to read, modify, and write back.
Unless you actually use writes for the whole process, writing the whole value each time with the bit changes.
There are tons of examples. Enabling RX interrupt, setting direction in half duplex mode, clearing a flag... Try all are masking operations requiring a read first.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 09, 2021, 02:22:26 pm
Just spent hours on this one. I created a second project in Cube (yeah - a complicated job!) and copied the 1st one to it.

Then I did a trivial change to make it flash some LED, to make sure Cube was really building the 2nd project. Well, it did flash the LED but it was massively smaller!

Spot the difference:

Code size about 250k:

Code: [Select]

for (uint32_t i=0; i<0xffffffff; i++)
{
LED_On(KDE_LED3);
hang_around(200);
LED_Off(KDE_LED3);
hang_around(200);
}
main();

The 30k project:

Code: [Select]

for (;;)
{
LED_On(KDE_LED3);
hang_around(200);
LED_Off(KDE_LED3);
hang_around(200);
}
main();

Obviously, the compiler is realising main() is never entered so it chucks out everything after that.

Title: Re: GCC compiler optimisation
Post by: DavidAlfa on August 09, 2021, 02:47:21 pm
Yeah, you have that option in the linker settings, "remove unused sections".
Title: Re: GCC compiler optimisation
Post by: ataradov on August 09, 2021, 04:14:16 pm
It does make me wonder whether these optimisations have any impact whatsoever on system performance.
They do. Remember, the compiler you are using is not specifically for embedded. It is the same compiler that works on PCs, and most those optimizations happen before target-specific code is generated.

Any optimization has impact on performance, it would not be optimization otherwise.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 09, 2021, 04:19:09 pm
Just spent hours on this one.
Why? When something is not clear. "arm-none-eabi-objdump -d file.elf". It would be trivially obvious when code is missing.

Also recent versions of GCC in some cases when they can detect undefined behaviour, will replace the whole chunk of code that relies on the UB with a single UDF instruction.

I ran into it when one control path accidentally used a local pointer that was not initialized. The whole function got replaced with "UDF" and it was awesome because it caused an immediate exception rather than some random exception in the future.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 09, 2021, 04:58:06 pm
Would something like this get optimised (a read of CPU FLASH)

       uint8_t* p = (uint8_t*) 0x08000000;
      ch=*p++;

on the assumption that the compiler has not seen any code writing to that location? That would be bizzare, surely?

The above piece of code *alone* won't be optimized out as long as the given compiler defines integer to pointer conversion. Indeed, according to the std, this is implementation-defined. Whereas it's "defined" on most implementations around (and in particular compilers targetting embedded targets), this is not in itself portable (but common sense will tell you this anyway), and may not even yield any output code on some particular implementation.

That set aside, as in your case, implementation is surely defined, let's take a look at what happens if you omit the volatile qualifier.
If you read the same location several times, the compiler may (most compilers with optimizations enabled will) only access the location ONCE, and then re-use the same read value at each later occurence (as long as it can be statically inferred). This is particularly problematic when reading typical peripheral "registers", as for instance, reading a certain register in a loop waiting for some flag to toggle WILL not do what you intend if the pointer was not qualified volatile.

Again for the small piece of code you posted, there is possibly missing context for explaining how it will effectively be compiled. In isolation, it can't be optimized out. A pointer dereference will be honored. At least the first time it appears in code flow. It's successive accesses that may get optimized out.

volatile ensures that all accesses to a given variable will be honored in the order they appear. That doesn't just happen with pointer dereference either. It happens with any object qualified volatile.
A small example:

Code: [Select]
int Test1(volatile int n)
{
        return n * n;
}

int Test2(int n)
{
        return n * n;
}

Latest GCC, with -O3:
* In Test1, n will be copied on the stack and accessed twice. (Which admittedly is pretty weird.)
* In Test2, it's not the case.

x86_64 code:
Code: [Select]
Test1:
.LFB1:
.cfi_startproc
movl %edi, -4(%rsp)
movl -4(%rsp), %eax
movl -4(%rsp), %edx
imull %edx, %eax
ret
.cfi_endproc
.LFE1:
.size Test1, .-Test1
.p2align 4
.globl Test2
.type Test2, @function
Test2:
.LFB2:
.cfi_startproc
movl %edi, %eax
imull %edi, %eax
ret
.cfi_endproc
.LFE2:
.size Test2, .-Test2

Use "volatile" when a given object may be modified *outside* the scope of the current compilation unit (so, without the compiler being able to know it can be modified.)
When you do know it's not the case, don't use volatile, as it can yield pretty inefficient code as can be illustrated above.
Title: Re: GCC compiler optimisation
Post by: Siwastaja on August 09, 2021, 07:05:38 pm
You usually set up the registers, then you set the enable bit.
So you need to read, modify, and write back.

But obviously you don't need RMW for this, two writes are enough. I do that all the time, ST libraries do that as well. First write are the flags except enable bit, second write is the same but now with the enable bit set as well. Two compile-time constant values, usually! Easy for the compiler to optimize. ... except the register is qualified volatile preventing this optimization if you do |= or &= on the register directly.


This pattern is best, and actually somewhat used in ST's library code, surprisingly:

Code: [Select]
tmp = flags;
(volatile) peripheral = tmp;
tmp |= enable;
(volatile) peripheral = tmp;

Just two writes to the peripheral. Compiler is free to optimize the the tmp since it's not volatile, using two compile time constant loads for example.
Title: Re: GCC compiler optimisation
Post by: gf on August 09, 2021, 07:35:52 pm
A pointer dereference will be honored. At least the first time it appears in code flow. It's successive accesses that may get optimized out.

A pointer dereference must be honored of course, but only in the sense of "as if". The resulting code does not need to do any memory fetch. See lines 7 and 16 here (https://godbolt.org/z/T9x1KqcPP).

It is rather about values, and common sub-expression elimination. If the value of a sub-expression *p is already known at a particular point in the code flow, then it does not need to be re-evaluated (unless *p is volatile, or unless the memory location might have been invalidated in the meantime, according to the aliasing rules).
Title: Re: GCC compiler optimisation
Post by: peter-h on August 09, 2021, 09:45:59 pm
Interesting...

Should I use

volatile uint8_t* p = (uint8_t*) 0x08000000;

or

uint8_t* p = (volatile uint8_t*) 0x08000000;

The code does run BTW, with the compiler default setting (not sure what it is in Cube, out of the box), and correctly because it is correctly verifying the FLASH content.

Elsewhere I am using e.g. this code to check that a section of FLASH has been erased

Code: [Select]
// Check it is all FFs
for (address=0x080e0000; address<0x080fffff; address+=4)
{
data=*(volatile uint32_t*)address;
if ( data != 0xffffffff )
error++;
}

This is probably wrong too (but it works):

(https://peter-ftp.co.uk/screenshots/202108093512715422.jpg)

Where should the "volatile" go? Is it enough to put it in the initial declaration of addr i.e.

volatile uint32_t addr=0x08000000;

I can't understand how a compiler could optimise out "addr" however, given that the address being read is "obviously" continually changing (incremented by 512 each time).
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 09, 2021, 10:02:27 pm
Should I use

volatile uint8_t* p = (uint8_t*) 0x08000000;

or

uint8_t* p = (volatile uint8_t*) 0x08000000;

The second one is not correct. The compiler should give you a warning about it. You're assigning a pointer to volatile to a pointer to non-volatile. As the compiler will tell you, 'p' will just lose the volatile qualification.

The first one is correct.
I tend to do this myself instead, because to me it looks more consistent:

volatile uint8_t* p = (volatile uint8_t*) 0x08000000;

But your version (1) is correct. 'p' should be qualified volatile. The constant pointer on the right hand doesn't need itself to be volatile, so my version is probably unnecessarily verbose. The conversion will be implicit.
Title: Re: GCC compiler optimisation
Post by: gf on August 09, 2021, 10:32:43 pm
But your version (1) is correct. 'p' should be qualified volatile.

More precisely, not the variable p is qualified volatile here, but
Code: [Select]
volatile uint8_t *p;
declares p as (non-volatile) pointer variable, pointing to a volatile memory location. But yes, this is what peter-h wants for this particuar case.

If the pointer variable itself should be volatile, too, then the declaration should look like
Code: [Select]
volatile uint8_t* volatile p;
=> volatile pointer variable, pointing to a avolatile memory location.
Title: Re: GCC compiler optimisation
Post by: gf on August 09, 2021, 10:53:14 pm
I can't understand how a compiler could optimise out "addr" however, given that the address being read is "obviously" continually changing (incremented by 512 each time).

The compiler could certainly call the function AT45dbxx_WritePage() with the constant (uint8_t*)0x08000000 as first argument, and don't allocate any register or stack frame slot for the variable "addr".


Sorry, too late in the night :=\ Overlooked the increment.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 10, 2021, 01:37:32 am
Where should the "volatile" go? Is it enough to put it in the initial declaration of addr i.e.

volatile uint32_t addr=0x08000000;

I can't understand how a compiler could optimise out "addr" however, given that the address being read is "obviously" continually changing (incremented by 512 each time).

Yeah. Your question shows that understanding "volatile" is much trickier than it looks for many people.
Possibly you're confusing the "address" and the read operation itself. Possibly same with the pointer.

For the code in your screenshot, there's still one potential pitfall. But it's not in the piece of code you're showing us. It's with the AT45dbxx_WritePage() function you're calling, and it's possibly yet another related problem that we haven't touched quite yet.

Assuming this function was defined within the same compilation unit (which typically means, either within the same source file, or in a source file itself included in this source file), the compiler could decide, depending on how the function is written, that it may have NO effect whatsoever when called with a pointer which points to no known object, and thus decide to optimize the call out entirely. The possible side-effect would be that 'addr' itself would never get incremented, unless you use the 'addr' value after the 'for' loop. If you don't, the optimizer may not even actually generate any code for addr. It would for page though, because page is used in an 'if' condition which itself has effects. Assuming the KDE_LED_xx() functions themselves are not optimized out. Are you starting to like it?

An example of function, used in place of AT45dbxx_WritePage(), that could trigger this behavior, is memset() or memcpy(). If you called memset() with a pointer to something converted from an int (some 'address' to something that's not known by the compiler), the call to memset() could be optimized out. Reason is that the compiler may assume that this call would have no effect.

One way to circumvent this is to write your own memset(), or memcpy() function. With a pointer to volatile as parameter. The std ones don't have the volatile qualifier in their parameters. This is something that can bite your ass here.

So for instance, for an "always has an effect" version of memset(), regardless of what you call it with, you would need to write your own.
The original std one has the following prototype: "void *memset(void *str, int c, size_t n)".
Yours should have this one: "void *my_memset(volatile void *str, int c, size_t n)".

A common trap is believing that merely casting the passed pointer to memset() with a volatile qualifier will do the trick, such as: "memset((volatile void *) p, ..., ...)". It won't, for the reason explained above about your pointer casts. This is actually an example I remember was discussed in another thread on this forum.

Title: Re: GCC compiler optimisation
Post by: ejeffrey on August 10, 2021, 03:18:45 am
Interesting...

It does make me wonder whether these optimisations have any impact whatsoever on system performance. I have decades' experience of assembler, and all the tricks people used to do (including self modifying code, which I avoided), so I understand this stuff at the machine level. And in most systems some 1% of the code is speed critical, and one generally gains far more there by sitting down and thinking about doing that job differently, than by rewriting it in the slickest assembler possible.

But that's not how optimizing compilers work.  They don't try to figure out that certain areas are hot and then deliberately produce unoptimized code for everything else because it isn't performance critical in the application.  They just try to produce high performance code everywhere.  To answer you question, yes all those optimizations make a difference, but obviously the difference is only large when the code in question is heavily used.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 10, 2021, 06:20:00 am
This stuff, especially SiliconWizard's post above, is unbelievable. If "addr" is a problem, reading incrementing CPU FLASH addresses, it seems almost impossible to write working C code unless you are a total expert at what could go mysteriously wrong.

I think I will stick to this compiler version and its compiler options for ever, because right now everything is working :)

My actuarial life expectancy is about 20 years so this is a viable strategy :) I still run some software from ~1995 (in a winXP VM) so I have a solution...

I have often looked at the assembler generated, when single stepping, and it looks perfectly reasonable. There will not be any perf gain obtained by removing little bits of it. To make code run faster, one needs to think about critical portions, which are usually tiny.

BTW, I don't actually understand the syntax of
volatile uint8_t* p = (uint8_t*) 0x08000000;
I got it out of the ST libs, and it seems to work :) I don't use pointers in my own code.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 10, 2021, 06:38:56 am
It is a good idea to understand the syntax of the language you are using.

The syntax of this is easy. "p" is a pointer to a volatile location. The address of that location is 0x08000000. The expression in the right does not need to be volatile because it is never directly used to access the memory, it is only used to initialize the pointer.

A situation where volatile is necessary:

int a = ((volatile uint8_t*)0x08000000)[100]; // read 100th byte from the start of the flash. In this case there is no explicit pointer is created, so the value of the expression is used directly.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 10, 2021, 07:11:38 am
I usually understand this when it is explained, but it escapes soon :)

In asm one did this all the time but never thought about it in terms of a formal syntax.

Anyway, I looked around Cube and found that currently I am running with no optimisation:

(https://peter-ftp.co.uk/screenshots/202108104112720908.jpg)

Would it be correct that with no optimisation in GCC, "volatile" is not needed?
Title: Re: GCC compiler optimisation
Post by: newbrain on August 10, 2021, 09:41:58 am
Would it be correct that with no optimisation in GCC, "volatile" is not needed?
If some code working or not (not taking performance into account) depends on optimizations being enabled or not, that code is 99% wrong.

Volatile semantic (and atomic etc.) needed or not needed cannot depend on optimization.
A compiler is perfectly enabled to optimize or not your code regardless of its command line options.
That -O0 does what you expect the abstract machine defined in the standard would do, is just an implementation accident.

It is a good idea to understand the syntax of the language you are using.

This, QFT.
And the semantic too.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 10, 2021, 12:54:03 pm
"Volatile semantic (and atomic etc.) needed or not needed cannot depend on optimization."

I know one cannot dispose of a complex topic briefly, but I find it incredible that C is so full of traps. I have seen enough of it over many years to know that basically no program would run.

Can someone demonstrate GCC optimising out a straight read of 0x08000000, using the simplest syntax which I understand of e.g.

 uint32_t address = 0x08000000;
 uint8_t buffer[1000];
 memcpy(buffer,(char*)address,1000);  // memcpy(buffer,address,1000); also works but you get a compiler warning

especially if further down there is

 address+=512;

With writing to that address, that to me looks more dodgy, but why really? It is not a memory variable, which you could discard if you don't see it read later. I do get compiler warnings if I do say

 address=fred+3;

and if address is never later accessed, and if the compiler warns then it can safely discard that load (because it did warn about it).

If you have another RTOS thread picking up address then you have to either make it volatile, or assign something (harmless) from it so it appears to be used.

I have been using global variables for simple inter-thread comms and they never produced a warning in the code writing them, for some reason. I suspect there is a "scope" involved, and perhaps if you have

{
 uint32_t address=36;
 ...
 ...
}

that will warn, because in any case address is not valid outside of the {} block and if it doesn't get picked up within that block, that is a candidate for a warning and probably removal. But if you have declared a static at the start of a .c file

 uint32_t address;

and write to it inside some block then I see no warnings and the code does work. Warnings occur only if the variable is never referenced in the file. Hence I think there is a "scope" involved, outside which it doesn't care. If this wasn't the case then IMHO a lot of programs would never work. And from what I can see, at least with -O0, that scope is just the current block, for variables defined in that block. I have never seen warnings on statics (declared, but never referenced), and the algorithm for picking up statics which were "written to but not read afterwards" would be pretty interesting.

Another thing is that removing such code, without a warning, is unlikely to produce any performance or size gain, because after all the coder intended it to do "something", so all you have achieved is a broken program.
Title: Re: GCC compiler optimisation
Post by: gf on August 10, 2021, 02:46:57 pm
Can someone demonstrate GCC optimising out a straight read of 0x08000000, using the simplest syntax which I understand of e.g.

Yes: https://godbolt.org/z/he6avac4E

Edit: And clang even eliminates a memcpy() from a local buffer[] array to 0x08000000: https://godbolt.org/z/86KbWo93z
(-> copying undefined values from buffer[] to 0x08000000 is obviously not considered better than not copying anything)

And if buffer[] is static (i.e. zero-initialized), then clang even replaces memcpy() by memset(), and eliminates buffer[]: https://godbolt.org/z/8Es7xhsr3
Title: Re: GCC compiler optimisation
Post by: newbrain on August 10, 2021, 02:57:25 pm
Can someone demonstrate GCC optimising out a straight read of 0x08000000, using the simplest syntax which I understand of e.g.

 uint32_t address = 0x08000000;
 uint8_t buffer[1000];
 memcpy(buffer,(char*)address,1000);  // memcpy(buffer,address,1000); also works but you get a compiler warning

especially if further down there is

 address+=512;
Easy! (and ninjaed by gf...)
In this example (https://godbolt.org/z/8zYPc9nxo) on godbolt (always be praised!) the whole shebang is thrown away starting with -O2, and an empty loop is produced with -O1.
How is this not expected? There are, standing the declarations, no observable side effects of that code.

Then, some other random notes:
Quote
If you have another RTOS thread picking up address then you have to either make it volatile, or assign something (harmless) from it so it appears to be used.
No. Do things in the right way - if you rely on things that are not guaranteed by the standard, you will be bitten.
In this case volatile (though, in general, global variables are not very good inter-thread communication primitives...).

Quote
But if you have declared a static at the start of a .c file

 uint32_t address;

and write to it inside some block then I see no warnings and the code does work.
Be careful not to conflate scope, duration and linkage properties of an object.
That variable has file scope, static duration, but it has by default external linkage, meaning the compiler cannot know whether some other translation unit will reference the same object.
Were it declared static (so making that internal linkage), more optimizations are possible.
Yes it is confusing. The "static" storage class specifier affect both linkage and duration.

Quote
Another thing is that removing such code, without a warning, is unlikely to produce any performance or size gain, because after all the coder intended it to do "something", so all you have achieved is a broken program.
The compilers (and the standard) care not what a programmer intends.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 10, 2021, 03:45:32 pm
uint32_t address = 0x08000000;
 uint8_t buffer[1000];
 memcpy(buffer,(char*)address,1000);  // memcpy(buffer,address,1000); also works but you get a compiler warning

This specific example is perfectly fine, memcpy() will discard volatile anyway. And if memcpy() is an actual library function, not a compiler intrinsic, then it is 100% guaranteed to work.

The examples of it being optimized posted earlier are because buffer is never used for anything. As soon as you use the data in the buffer, a complete code will be generated.
Title: Re: GCC compiler optimisation
Post by: Siwastaja on August 10, 2021, 03:46:51 pm
For someone who understands how the CPU works and has written assembly for decades, C can be surprisingly difficult. In assembly, you directly enter the correct instruction, which is almost always obvious to you as the programmer. In C, there are two cases, the simpler "high-level" case where it doesn't matter, but when the right memory access pattern does matter, then you need to "guide" C through its type system, using pointers, casts, and possibly the volatile qualifier in a way that is different than just writing the right instruction directly.

But it's not impossible to learn, even at old age I'd guess. The rules are simple, it's just a matter of changing the perspective.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 10, 2021, 04:13:21 pm
ataradov - yes, so those are not great examples because that code fragment does nothing. This

Code: [Select]
int f()
{
    char buffer[512];
    memcpy(buffer, (char*)0x08000000, 512);
    char fred=buffer[0];
    return (int) fred;
}

compiles into real code in all cases.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 10, 2021, 04:23:34 pm
ataradov - yes, so those are not great examples because that code fragment does nothing. This
....
compiles into real code in all cases.
And it is expected that it would work. Compilers are not that smart with optimizations. But there is a difference between passing pointers to external functions and actions of the compiler itself. The functions are opaque for the compiler in most cases. memcpy() is not the best example, as it is a compiler intrinsic in a lot of cases, not an actual function, so more optimizations are possible.
Title: Re: GCC compiler optimisation
Post by: langwadt on August 10, 2021, 04:28:21 pm
ataradov - yes, so those are not great examples because that code fragment does nothing. This


sometimes the compiler is smart enough to see that the code does nothing (or something much simpler) even though you can't see it.

the optimizer makes the code do efficiently what you tell it to do, not necessarily how you tell it to do it
Title: Re: GCC compiler optimisation
Post by: gf on August 10, 2021, 04:29:12 pm
Code: [Select]
int f()
{
    char buffer[512];
    memcpy(buffer, (char*)0x08000000, 512);
    char fred=buffer[0];
    return (int) fred;
}

compiles into real code in all cases.

Clang still does not copy all 512 bytes to buffer[], but fetches only the first byte from 0x08000000, which is actually used at the end: https://godbolt.org/z/jGrv1vYj7
Title: Re: GCC compiler optimisation
Post by: peter-h on August 10, 2021, 04:42:06 pm
"sometimes the compiler is smart enough to see that the code does nothing (or something much simpler) even though you can't see it."

OK, but one would rarely write code which actually does nothing. Sometimes... one sets up a variable which is unused but that's rare.

"Clang still does not copy all 512 bytes to buffer[], but fetches only the first byte from 0x08000000, which is actually used at the end: https://godbolt.org/z/jGrv1vYj7"

That is hilarious!

In this case

Code: [Select]
int f()
{
    char buffer[512];
    memcpy(buffer, (char*)0x08000000, 512);
    char fred=buffer[0];
    fred+=buffer[511];
    return (int) fred;
}

it does just two reads; it is basically optimising out memcpy(), and doing it correctly, so the code would run correctly.

The time it would break is if the "thing" at 0x08000000 was something which was expecting to see 512 read cycles, but I would not use memcpy to achieve 512 contiguous read cycles because it is known that these functions do a load of optimisations e.g. on a 32F they would copy 32 bits at a time and then fix up the ends with byte reads (or some such). So yes this is a good example!

Title: Re: GCC compiler optimisation
Post by: gf on August 10, 2021, 05:04:38 pm
A (non-volatile) memory fetch or memory store is not considered a visible side effect, therefore the exact memory access pattern does not need to be preserved by the optimizer.
OTOH, a volatile memory fetch or memory store is by definition considered a visible side effect.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 10, 2021, 05:16:32 pm
So yes this is a good example!
And just to be clear, the only reason it works is because memcpy() is not a real function anymore, it is just an indication of intent to the compiler that is handled internally.  Many standard functions are handled this way. If you call printf() with just the string, it will be substituted with puts().

If you substitute your own version that just copies things in a loop, then it will not be optimized this way.

Which is why you need to have your own implementation for ordered access anyway. Standard functions do not guarantee order of writes or reads.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 10, 2021, 05:27:49 pm
Like this :)

Code: [Select]
eeread(0, buf);  // read 512 bytes into a buffer, from a serial EEPROM
    uint32_t fsize = buf[4]|(buf[5]<<8)|(buf[6]<<16)|(buf[7]<<24);

although I am sure there is a slicker way, e.g. overlaying a packed struct onto buf, but that would assume endianess. And with the 32 bit barrel shifter the above will still be really fast.

This is quite a learning experience! But fortunately I think all my code will work fine.
Title: Re: GCC compiler optimisation
Post by: gf on August 10, 2021, 05:39:34 pm
If you substitute your own version that just copies things in a loop, then it will not be optimized this way.

Unbelievable - even then it is optimized out :-DD https://godbolt.org/z/eE9761jns
(clang obvioulsy still recognizes what the copy() does - at least if it can be inlined)
Title: Re: GCC compiler optimisation
Post by: peter-h on August 10, 2021, 05:46:18 pm
What about 32F4 ARM GCC - I seem to be on 9.3.1.

I would expect x86 compilers to be clever.

GCC produces

Code: [Select]
f:
        sub     sp, sp, #512
        mov     x2, -513
        add     x0, sp, 512
        mov     x3, 512
        movk    x2, 0xf7ff, lsl 16
        movk    x3, 0x800, lsl 16
        add     x2, x0, x2
        mov     x0, 134217728
.L2:
        mov     x1, x0
        add     x0, x0, 1
        cmp     x0, x3
        ldrb    w1, [x1]
        strb    w1, [x0, x2]
        bne     .L2
        ldrb    w0, [sp]
        add     sp, sp, 512
        ret

Not as clever!

With -O0, it does it literally

Code: [Select]
copy:
        sub     sp, sp, #48
        str     x0, [sp, 24]
        str     x1, [sp, 16]
        str     x2, [sp, 8]
        ldr     x1, [sp, 16]
        ldr     x0, [sp, 8]
        add     x0, x1, x0
        str     x0, [sp, 40]
        b       .L2
.L3:
        ldr     x1, [sp, 16]
        add     x0, x1, 1
        str     x0, [sp, 16]
        ldr     x0, [sp, 24]
        add     x2, x0, 1
        str     x2, [sp, 24]
        ldrb    w1, [x1]
        strb    w1, [x0]
.L2:
        ldr     x1, [sp, 16]
        ldr     x0, [sp, 40]
        cmp     x1, x0
        bne     .L3
        nop
        nop
        add     sp, sp, 48
        ret
f:
        sub     sp, sp, #544
        stp     x29, x30, [sp]
        mov     x29, sp
        add     x0, sp, 24
        mov     x2, 512
        mov     x1, 134217728
        bl      copy
        ldrb    w0, [sp, 24]
        strb    w0, [sp, 543]
        ldrb    w0, [sp, 543]
        ldp     x29, x30, [sp]
        add     sp, sp, 544
        ret

In most applications the extra time will not be of relevance but the doubling of code size may well be.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 10, 2021, 06:11:04 pm
The rules are actually pretty simple. What makes them hard to apply for us humans is that we inherently have a very hard time thinking that any single statement of code we write could be actually useless. That's probably because we are just so full of ourselves. ;D

One thing that this kind of topic teaches you is that C is definitely NOTHING LIKE A PORTABLE ASSEMBLER, despite what many uninformed people keep saying. C statements are absolutely not guaranteed to be translated verbatim to machine code.

This does not mean that it's impossible to write "working C code". Most issues related to the kind of optimizations we're talking about in this thread deal either with objects that the compiler doesn't know about, or from expectations related to the above paragraph otherwise. For instance, expecting some statement to actually yield verbatim machine code when said statement doesn't actually change any result from the program's execution.

A typical example, often discussed in forums, but not in this thread, is the famous delay loop. You'll just write an empty 'for' loop in hopes it will actually yield the machine code executing for some time. Optimizers will just prune such loops. Because they have NO effect from a "functional" POV. They do not change the result of anything.

Code: [Select]
for (int i = 0; i < xxx; i++) {}
The way to force this generating an actual loop is either to declare the counter variable volatile, or do something in the loop body that the compiler can't prune. So this would be either:

Code: [Select]
for (volatile int i = 0; i < xxx; i++) {}
or:

Code: [Select]
for (int i = 0; i < xxx; i++) { Nop(); }
Nop() here being for instance some macro that actually is some assembly code (whatever it is, it must be something that the compiler can't assume having no effect). Drawback of the second version is that it's not portable. Benefit is that the loop will be usually more "efficient", because the loop counter here can be put in a register, whereas in the first form, the 'i' counter will usually be put on the stack, and a read access, incrementation, and write access will occur at each iteration. That's usually what happens. Again, there is no guarantee that a given compiler will compiler this in any specific way, but there's a guarantee that either of the two above forms will be compiled as actual loops.

Dealing with "outside objects" (such as any register/buffer accessed through pointers to absolute addresses) is indeed not that simple in C, while C is said to be a very low-level language good at this. Many developers will actually use C only for embedded dev these days... for which this kind of issues are ubiquitous. It doesn't take being an "expert", but I admit it takes more knowledge than is usually thought. C is again too often thought about as a very simple language. It's not quite.

Now what could be nice, in order to help developers with this, would be that compilers would actually give a list of any piece of code that would have been pruned during optimization. Problem with this is  that in a typical program, the list is likely to be pretty large, and it would take you a lot of time to just go through everything in hopes of catching something getting pruned that you actually want to execute...
Title: Re: GCC compiler optimisation
Post by: peter-h on August 10, 2021, 06:19:12 pm
One could still issue warnings in some cases - like that ludicrous removal of most of my program above :)

I have just tried a bigger piece of real code. This one is not right; it was hacked since this compiler doesn't know uint32_t and some other stuff...

Code: [Select]

            void L_HAL_FLASH_Unlock(void);
            void L_HAL_FLASH_Lock(void);
            void L_FLASH_Program_Word(int addr, int data);

        static char buffer1[32768];
           void test (void)
           {
           
            int page,pagebase;
            int error=0;
            int cpubase=0x0800000;
            int AT45dbxx_ReadPage(char* buf, int count1, int count2);
            int cust_blocks=444;

            for (int block32k=1; block32k<cust_blocks; block32k++) // 1-31
  {

  // Read each 32k block into buffer1
  int buffer1idx=0;
  for ( int page=pagebase; page<(pagebase+64); page++ )
  {
  AT45dbxx_ReadPage(&buffer1[buffer1idx],512,page*512);
  buffer1idx+=512;
  }

  // Program 32k block
  L_HAL_FLASH_Unlock();
  for (int i=0; i<(32*1024); i+=4)
  {
  int data=buffer1[i]|(buffer1[i+1]<<8)|(buffer1[i+2]<<16)|(buffer1[i+3]<<24);
  L_FLASH_Program_Word(i+cpubase, data);
  }
  L_HAL_FLASH_Lock();

  // Verify 32k block against buffer1
  for (int i=0; i<(32*1024); i+=4)
  {
  int data=buffer1[i]|(buffer1[i+1]<<8)|(buffer1[i+2]<<16)|(buffer1[i+3]<<24);
  if ( (*(volatile int*) (i+cpubase)) != data ) error++;

}

  cpubase+=(32*1024);
  pagebase+=64;
  block32k++;

  }
            }


With -O0 I get 161 lines. With -O1 I get 67, and after that not much changes. It is interesting because most of the difference is just using different instructions. I don't see it doing anything cunning.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 10, 2021, 07:04:06 pm
Unbelievable - even then it is optimized out :-DD https://godbolt.org/z/eE9761jns
(clang obvioulsy still recognizes what the copy() does - at least if it can be inlined)
Wow. This is pretty cool.

So yeah, volatiles where needed is a must. Anything else is just a trouble waiting to happen.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 10, 2021, 08:04:20 pm
That x86 compiler must be building a table for every element of every array and every variable, to keep track of each element accessed by the program.

So if e.g. you had

uint8_t buf[512];

for (uint32_t i=0;i<256;i++)
{
 buf[ i ]='a' ;
}

uint8_t fred=6;
i++;
buf[i ]=fred;

and if buf[257-511] was never used after that, it would throw away that code. Well, correctly so, because it would not do anything.

It must be almost "running" the code to do this.

I knew a guy many years ago who wrote a bunch of C compilers. He spent a lot of time looking for special cases e.g. integer x 10 is much faster (on the old CPUs) if you do x2, x8, and add. But you can always do that, 100% safely. I don't recall anybody using "volatile" in the 1980s. C just did what you expected :) I supervised some quite big developments done in IAR C; I did asm portions to speed them up (in some cases by 100x to 1000x) relative to an admittedly dumb use of sscanf.
Title: Re: GCC compiler optimisation
Post by: gf on August 10, 2021, 09:10:06 pm
What about 32F4 ARM GCC

ARM Cortex M4 clang: https://godbolt.org/z/4Eh1faYr9
ARM Cortex M4 gcc:    https://godbolt.org/z/KdvhdW4zM
x64 clang:                    https://godbolt.org/z/rd1MTx5bG
x64 gcc:                       https://godbolt.org/z/39hdTE8Ph

clang optimizes both, memcpy() and copy(), for both processors.
gcc does even a better job for ARM than for x64, in this particular case.
But you can't not generalize that to other code. It really depends.
Title: Re: GCC compiler optimisation
Post by: westfw on August 10, 2021, 09:35:04 pm
Quote
because memcpy() is not a real function anymore, it is just an indication of intent to the compiler that is handled internally.  Many standard functions are handled this way.
Is there a command-line switch to stop this interpretation of "standard" functions?  (globally, or on a per-function basis?)

I find it annoying; one of the "nice" things about C was that functions (including all of IO) were functions, and not behavior hard-wired into the language. (also one of the reasons that C is comparatively easy to port to new platforms.)

For a language often derided as "a high level assembly language", the compiler folk seem very keen on doing things to it that will annoy people who are used to assembly language.  :-(

For a less controversial example, consider something like:
Quote
    *flashPtr = x;  // load up flash write buffer

    // Manipulate NVMCTL to actually write the flash buffer to flash
    //  :

    if (*flashPtr != x) {   // verify
         // flash write error!
    }
if *flashPtr is not "volatile", even a relatively dumb compiler would have no reason to re-read it; it has not way to know that manipulating the NVMCTL stuff might change the contents of memory...
Title: Re: GCC compiler optimisation
Post by: peter-h on August 10, 2021, 09:58:02 pm
"gcc does even a better job for ARM than for x64, in this particular case."

That's v10, which seems quite different to v9.x.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 10, 2021, 10:04:00 pm
Is there a command-line switch to stop this interpretation of "standard" functions?  (globally, or on a per-function basis?)
Sure. "-fno-builtin" for all of them, and then there are flags like "-fno-builtin-memcpy".

Title: Re: GCC compiler optimisation
Post by: emece67 on August 10, 2021, 10:34:53 pm
.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 10, 2021, 10:42:36 pm
Declaring any of to or from as pointer to volatile results in copy() compiled as a loop.
Yes, sure. But the fact that it can track this stuff across function calls is great.

If you add "-fno-builtin", it will also generate the full loop. So it recognizes the copy semantics in that function, and it has logic to optimize that specifically.

This optimization is not generic. Replacing the body of the loop with "*to++ = *from++ * 2" breaks it, since it is not longer a simple copy.
Title: Re: GCC compiler optimisation
Post by: ejeffrey on August 10, 2021, 11:34:49 pm
Quote
because memcpy() is not a real function anymore, it is just an indication of intent to the compiler that is handled internally.  Many standard functions are handled this way.
Is there a command-line switch to stop this interpretation of "standard" functions?  (globally, or on a per-function basis?)

I find it annoying; one of the "nice" things about C was that functions (including all of IO) were functions, and not behavior hard-wired into the language. (also one of the reasons that C is comparatively easy to port to new platforms.)

I think -fno-builtins will do it?  There are a number of built-in function and options to control their use, I don't remember all of them, mostly they are useful for writing the C library or other platform implementations.

It's not really true that memcpy isn't a real function.  There is absolutely a version in libc that will be called if necessary (for instance if using function pointers).  But the compiler has the option to replace it with equivalent code if possible, and that choice can be platform and context sensitive, such as if the compiler can prove alignment conditions or a compile time length.  It's not really any different from inline functions.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 10, 2021, 11:42:32 pm
It's not really true that memcpy isn't a real function.  There is absolutely a version in libc that will be called if necessary (for instance if using function pointers).  But the compiler has the option to replace it with equivalent code if possible, and that choice can be platform and context sensitive, such as if the compiler can prove alignment conditions or a compile time length.  It's not really any different from inline functions.

Indeed, the compiler can do similar optimizations with user-defined functions. But it can just do better with std functions because it knows what they are supposed to achieve exactly.

But any kind of pruning - it can absolutely do the same. Thus the importance, *when relevant*, to qualify pointer parameters with volatile.
Title: Re: GCC compiler optimisation
Post by: cfbsoftware on August 10, 2021, 11:58:05 pm
One thing that this kind of topic teaches you is that C is definitely NOTHING LIKE A PORTABLE ASSEMBLER, despite what many uninformed people keep saying. C statements are absolutely not guaranteed to be translated verbatim to machine code.
I would not interpret the statement that 'C is like a portable assembler' as 'C is a portable assembler'. I see it as simplistic way to describe that C is best-suited to write software that you might otherwise have to use assembler for.

Title: Re: GCC compiler optimisation
Post by: Bassman59 on August 11, 2021, 03:44:53 am
I don't recall anybody using "volatile" in the 1980s. C just did what you expected.

I'm in that category! I don't remember using "volatile" in Turbo C or really any personal-computer C programming. The first time I came across it was when I started using C for the 8051 back in, oh, 1997, I think. And the reason for "volatile" was for the usual reason: to tell the compile to not optimize access to global or file-scope variables that could be changed in an ISR.

And now that I think about it, I don't recall writing interrupt handlers in C for the PC under DOS and I know I never wrote any Windows programs that used interrupts. Of course embedded stuff on (what we now call) bare-metal 8051 and later for me 68k interrupts were necessary.
Title: Re: GCC compiler optimisation
Post by: ejeffrey on August 11, 2021, 04:04:42 am
I knew a guy many years ago who wrote a bunch of C compilers. He spent a lot of time looking for special cases e.g. integer x 10 is much faster (on the old CPUs) if you do x2, x8, and add. But you can always do that, 100% safely. I don't recall anybody using "volatile" in the 1980s.

A lot of people wrote incorrect code in the 80s, whats your point?  More to the point, ANSI C wasn't even fully standardized until 1989, until then you weren't targeting a standard you were targeting an implementation.  That is also one reason performance code was usually written in Fortran instead of C -- Fortran had much better optimizers available.

For user applications you rarely need volatile anyway.  You only need it or should use it for memory mapped IO and (carefully) for signal / interrupt handlers.  Non-kernel code written for UNIX platforms of the day wouldn't have needed volatile for much. DOS on x86 required user applications to do hardware access although x86 had IO instructions for IO port access to which none of this applies.

In the 80s and 90s inlining was not well standardized and generally only performed when specifically requested.  Link-time optimization across translation units basically didn't exist.  If you put your IO in functions in a library and didn't mark them inline, many of these optimizations could not kick in if the compiler can't look across function call barriers.

Quote
I supervised some quite big developments done in IAR C; I did asm portions to speed them up (in some cases by 100x to 1000x) relative to an admittedly dumb use of sscanf.

Re-writing in assembly because sscanf is too slow seems a bit overkill, but whatever.  Just rewriting in C would probably get you almost all the benefit.  sscanf and friends are extremely slow.  They are designed to be flexible and simple to use, not high performance.
Title: Re: GCC compiler optimisation
Post by: newbrain on August 11, 2021, 06:37:56 am
That is also one reason performance code was usually written in Fortran instead of C -- Fortran had much better optimizers available.
Just one nitpick:
Fortran had some advantages on C89 due to the different aliasing constraints that allowed for better optimizations.
This was (partially?) solved when C99 introduced the 'restrict' type qualifier.
In fact, memcpy signature now takes restricted pointers, so the behavior is undefined if the objects overlap: this was implicit (document) pre C99, but is now explicit part of the contract in the declaration.

C++ does not support restrict, though most compilers do as an extension.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 11, 2021, 02:37:12 pm
Could any of you experienced chaps suggest on which of these are worth enabling?

I have two configs: Debug (-O0 and max debug and Release (-O3 and no debug).

(https://peter-ftp.co.uk/screenshots/20210811243683315.jpg)

O3 is picking up some interesting warnings. One of these is

Code: [Select]
memcpy(inbuf,inbuf+1,INBUF_SEARCH_LEN); // Shuffle search area one left (its last byte is now garbage)
and I thought this is relevant in light of the above memcpy discussion. Memcpy is supposed to copy overlapping buffers correctly, no?? It's been running for months.

Inbuf is 1000 bytes long, and the search length is 7 bytes. It is a primitive way of testing the first 7 bytes for one of about a hundred different strings (to do with GPS satellite vehicle IDs etc). I then do

Code: [Select]
if ( memcmp(inbuf,pubx_00_header,7) == 0)
Title: Re: GCC compiler optimisation
Post by: oPossum on August 11, 2021, 02:48:38 pm
Memcpy is supposed to copy overlapping buffers correctly, no?? It's been running for months.

No!  memmove() allows overlapping buffers.
Title: Re: GCC compiler optimisation
Post by: newbrain on August 11, 2021, 02:52:25 pm
Could any of you experienced chaps suggest on which of these are worth enabling?

I have two configs: Debug (-O0 and max debug and Release (-O3 and no debug).
[...]
Memcpy is supposed to copy overlapping buffers correctly, no??

Inbuf is 1000 bytes long, and the search length is 7 bytes. It is a primitive way of testing the first 8 bytes for one of about a hundred different strings (to do with GPS satellite vehicle IDs etc). I then do

Code: [Select]
if ( memcmp(inbuf,pubx_00_header,7) == 0)
You might have noticed I'm a bit of a pedant - only with C, I swear!
So my code usually goes with -pedantic -Wextra -Wall -Wswitch-default -Wswitch-enum (plus -std=c11 usually).
Of course, these might be  relaxed for library (not mine) code.

No, memcpy has undefined behaviour if the objects overlap.
As previously noted, look at the signature (C99):
Code: [Select]
void* memcpy( void *restrict dest, const void *restrict src, size_t count );The 'restrict' type qualifier makes it clear that the compiler does not expect the memory pointed by the arguments to overlap.
So, it's UB if they do.
In C89, this was in the documentation - restrict (added in C99) makes that explicit.

To copy overlapping buffers, you need memmove (https://en.cppreference.com/w/c/string/byte/memmove).

As a habit, I use -Og for debug compiles: the code remains eminently debuggable, but a lot of pointelss memory/register shuffling and pushing/popping is removed.

EtA: -Wconversion is a bit chatty for my tastes. Helpful if you think you might have gotten some of them wrong.
Title: Re: GCC compiler optimisation
Post by: newbrain on August 11, 2021, 03:03:17 pm
It is a primitive way of testing the first 7 bytes for one of about a hundred different strings (to do with GPS satellite vehicle IDs etc). I then do
Wait, what was the problem with using another pointer (e.g. char *search_ptr = inbuf + 1;) instead of moving bytes around?

Code: [Select]
if ( memcmp(search_ptr, pubx_00_header, 7) == 0)
Title: Re: GCC compiler optimisation
Post by: peter-h on August 11, 2021, 03:37:33 pm
I googled on this, because I "asked somebody" before using memcpy for that, and just found this

87. WHAT IS THE DIFFERENCE BETWEEN MEMCPY() & MEMMOVE() FUNCTIONS IN C?
memcpy()  function is is used to copy a specified number of bytes from one memory to another.
memmove() function is used to copy a specified number of bytes from one memory to another or to overlap on same memory.
Difference between memmove() and memcpy() is, overlap can happen on memmove(). Whereas, memory overlap won’t happen in memcpy() and it should be done in non-destructive way.


It is bad grammar but it seems to be saying that if you want to shuffle a 7 byte buffer 1 place, you need to use memmove.

Re the shuffling, I was just being "simple" :) I am looking for a pattern in a data stream arriving on a serial port, so one can't just keep incrementing a pointer because when it gets to end of buffer minus 7, the compare string will run off the end. And anyway, upon match, one needs to fetch the rest. I know there are super slick ways of doing this stuff...

Another thing I am noticing. If one switches builds, one needs to do a Clean Project, otherwise it doesn't rebuild the whole project (or any of it actually) because it finds all the .o files already in place.

I am using -Og now, and I get 160k (down from 225k with -O0 and max debugs) in Debug build versus 179k in Release build with -O3. How is that possible?

Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 11, 2021, 04:14:45 pm
Could any of you experienced chaps suggest on which of these are worth enabling?

-Wall
-Wconversion

This last one, I talk about on a regular basis. It will just warn you about all dubious implicit conversions. Very useful, especially for embedded development.

-Wextra can get you a lot of extra noise with not a lot of added value, compared to the two above. But you can experiment and see for yourself.

In any case, I suggest reading this to better understand what those flags are all about: https://gcc.gnu.org/onlinedocs/gcc/Warning-Options.html
Title: Re: GCC compiler optimisation
Post by: ejeffrey on August 11, 2021, 04:22:29 pm
Quote from: peter-h
I am using -Og now, and I get 160k (down from 225k with -O0 and max debugs) in Debug build versus 179k in Release build with -O3. How is that possible?

Not sure what size you are reporting, but why do you think this is strange?  None of those options are deliberately optimizing size, so the variation in size is incidental.

-O0 will have a bunch of extraneous and redundant data movement operations so it will be big.  -Og, -O1, and -O2 will enable a bunch of optimizations, and many of those will involve removing redundant instructions or combining instructions leading to smaller code size.  The gcc man page for -O2 (you did read that, right?) says " performs nearly all supported optimizations that do not involve a space-speed tradeoff."  although if you actually look at the list of optimizations some of them do moderately increase size.  -O3 enables (in particular) a number of loop optimizations that can increase code size.  If you want the smallest code size you need -Os.
Quote from: newbrain
As a habit, I use -Og for debug compiles: the code remains eminently debuggable, but a lot of pointelss memory/register shuffling and pushing/popping is removed.

As you should, -O0 won't give you all the warnings you want.  At least in the past, -O0 disabled reachability analysis so it wouldn't warn you about unreachable code, computed values never used, or use of uninitialized variables.  I'm not sure what the current behavior is exactly but definitely the recommendation continues to be to use -Og for debug builds to get maximum diagnostics.  I think the only reason to use -O0 is to get maximum compilation speed.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 11, 2021, 07:26:23 pm
Very interesting.

-O2 gives me 160k, versus 179k with -O3.

I have switched the Debug build to -Og.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 11, 2021, 07:30:12 pm
With GCC at least, -O3 almost always yields larger code size. The main reason for this that I've seen is that the compiler will aggressively inline every function it can at -O3.
Up to you to determine whether -O3 is worth it performance-wise in your particular application. Sometimes it is, sometimes it's not.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 11, 2021, 09:07:35 pm
I wonder how this maps onto the 32F4 cache, which loads x bytes (?) from the FLASH at a time into zero-waitstate RAM, plus there is the instruction prefetch. That scheme is likely to run a loop (which fits into x bytes) a lot faster than unrolling the loop. Inlining functions probably is better because it would work with the prefetch.

Inlining functions is ok for small functions, obviously, but big ones will massively swell the code.

EDIT: I have just found that -Og has optimised out a huge amount of my startup code. The rest is basically running but I have e.g. a 512 byte buffer into which I read some eeprom data and then look at various values in that, and I think it stripped out all that code. It doesn't make sense to me why. Maybe putting 'volatile' in front of every variable might be worth a try. So I am back to -O0 and it works just fine.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 11, 2021, 11:35:54 pm
At this point putting an effort into understanding volatiles might be a better option.

If compiler strips out half your program, you are clearly doing something wrong. Why struggle? Even if you just want to brute force it, set seemingly related variables as volatile and see which ones really need to be volatile.

I would be very scared to work with a fragile code like this.

Also, compilers sometimes will strip significant portions of the code if it includes undefined behaviour, since one version of the UB code is no better than any other, so you might as well strip it.
Title: Re: GCC compiler optimisation
Post by: ejeffrey on August 12, 2021, 01:11:04 am
Another possibility is that the problem is not with volatile but with pointer casting abuse running afoul of the aliasing rules.  Only the hardware registers should really need to be volatile from what I can piece together.

For instance you have a block of memory that is zero initialized, then load it from EEPROM via a function that takes an integer rather than char * as a parameter, then do some computation the compiler may assume the data is still zero and optimize away everything that touches it. Peppering the code with volatile might make it work properly but isn't the actual bug.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 12, 2021, 05:42:51 am
I will investigate tonight but basically I am reading a 512 byte block from an eeprom, into a buffer on the stack. Then picking some boolean flags out of that, and based on those I do various other things.

That buffer is being incorrectly filled and is full of 0x55 which is the memory init value. The ports are all volatile-defined. ST use the __IO prefix which is #defined as volatile.

It looks like the compiler is seeing the content of the buffer as static values and deciding the booleans must be false and then removing a load of later tests on that basis.

Never seen anything like it... but it won't be hard to dig around, because when stepping you see half the variables showing as "optimised out".

Obviously I need to find out why that buffer read fails, but it is complicated code, from the ST SPI lib and the ST (or Adesto) 45DBxx lib.

Also stuff like

Code: [Select]
        uint8_t ssa_buf[512];
SSA_read(0, ssa_buf);
    uint32_t size1 = ssa_buf[4]|(ssa_buf[5]<<8)|(ssa_buf[6]<<16)|(ssa_buf[7]<<24);
        uint32_t size2 = ssa_buf[8]|(ssa_buf[9]<<8)|(ssa_buf[10]<<16)|(ssa_buf[11]<<24);

where the two uint32s are being removed. Previously they were removed if not referenced, which is normal. They are valid statements regardless of the content of ssa_buf.

My code is damn simple. No pointers :) I do pass some parms by address but very rarely and only if necessary to do the job (like the crc accumulator in a one byte at a time crc func).

I had a quick look and I see this utterly weird thing:

(https://peter-ftp.co.uk/screenshots/202108121812805806.jpg)

code starts off ok, but by the time I get to the bottom (the green-highlighted line) it is showing as 'optimised out'. How is that possible?

Title: Re: GCC compiler optimisation
Post by: newbrain on August 12, 2021, 06:19:05 am
Please note that when you see 'optimized out' while debugging, it only means that the generated code is not reserving a real variable in memory but it dos not mean it's doing away with the statements you have written.

Classic examples are temporary variables one might use to hold intermediate results in a calculation, quite often also loop indexes etc.

That said, I would really generalize ataradov's advice about learning the meaning (semantics) of C: I'm under the impression you've been trying to run before you could walk (and plagued by having to cope with Eclipse - barf)
On how to do this, unfortunately I cannot really say: having learned bad C from Schildt's books in a distant past, my salvation came only when I took the the (C89, at the time) standard and its rationale document and studied it cover to cover - but I understand that's not for everyone...

Sorry if I come out as patronizing, but the sheer number of posts does seem to indicate some basic (well, actually C  ;)) issue. Not that I don't enjoy the discussion, there is always room to learn and refine one's knowledge!
Title: Re: GCC compiler optimisation
Post by: ataradov on August 12, 2021, 06:36:01 am
Are you sure the code is actually removed? Eclipse may just be confused by the debug information. Those variables may have just ended up in registers, so no memory is needed for them. Technically this is indicated in the debug information, but ability of tools to interpret this information varies.

The easiest way to debug stuff like this is to "highlight" the section with nops and look in the disassembly:

Code: [Select]
asm("nop");
asm("nop");
asm("nop");
uint32_t size1 = ssa_buf[4]|(ssa_buf[5]<<8)|(ssa_buf[6]<<16)|(ssa_buf[7]<<24);
uint32_t size2 = ssa_buf[8]|(ssa_buf[9]<<8)|(ssa_buf[10]<<16)|(ssa_buf[11]<<24);
asm("nop");
asm("nop");
asm("nop");

The code for the expressions would be placed between the nops. You can look at the code and see what it is doing exactly and you will see where those values are stored.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 12, 2021, 07:10:45 am
OK, yes, disassembly shows the code is there; 'code' is held in a register. I am very familiar with asm :)

Amusingly, I think this is where one issue is: a delay func for use where there is no timer tick

Code: [Select]
// Hang around for delay in ms. Approximate but doesn't need interrupts etc working.

static void hang_around(uint32_t delay)
{

uint32_t fred = 17000*delay;

while (fred>0)
{
fred--;
}

}

Putting volatile there makes a lot of stuff work :)

newbrain - the reason I post a lot is because I have nobody else I can ask. I am working more or less alone. I have nobody who is available as a "C consultant". I have actually tried to set something like that up (via my little business) but nobody is interested. I paid one guy £500 to configure a server with a PHP prog running on an existing server, and he charges 50/hr for extra work, which is fine, but wanted 500/month for ongoing "support" which is way too much for what will be needed (maybe 1hr/month). Then I paid another guy ~10k to write a PHP site, from scratch, to a spec, which worked out well and he charges similarly for ongoing, but doesn't do embedded. These are all in the former Soviet Bloc, of course :) As it happens, I am looking for someone to do well defined portions of this current project too but it is too early to do that since the API is not all done and documented. I do a lot of googling of course but sometimes a focused Q on a good forum is a good way, and the vast majority of stuff online is garbage... all the way to code examples which could not have ever worked. I have always used forums, for various topics, not just electronics, and run one myself. But I am learning. It is a steep curve in places. I am documenting everything too, with lots of comments in the code (most progs don't comment C at all) :)

I appreciate everyone's help, and hope that some others quietly reading this stuff might find it useful.

Title: Re: GCC compiler optimisation
Post by: ataradov on August 12, 2021, 07:16:15 am
A much better version of a blocking delay:

Code: [Select]
__attribute__((noinline, section(".ramfunc")))
void delay_cycles(uint32_t cycles)
{
  cycles /= 4;

  asm volatile (
    "1: sub %[cycles], %[cycles], #1 \n"
    "   nop \n"
    "   bne 1b \n"
    : [cycles] "+l"(cycles)
  );
}
The instruction will be either sub or subs depending on the core type. The code assumes no wait states, so placed in SRAM. But you can obviously leave it in the flash, just adjust the division factor.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 12, 2021, 07:23:57 am
Is that better because the asm will not get modified by optimisation?

I guess fred will get put into a register...
Title: Re: GCC compiler optimisation
Post by: ataradov on August 12, 2021, 07:31:42 am
Is that better because the asm will not get modified by optimisation?
It is better because it is predictable and does not depend on the compiler.

I guess fred will get put into a register...
but the surrounding code can be anything and you will need to figure out the constant for different cases.
Title: Re: GCC compiler optimisation
Post by: newbrain on August 12, 2021, 07:39:47 am
OK, yes, disassembly shows the code is there; 'code' is held in a register. I am very familiar with asm :)

Amusingly, I think this is where one issue is: a delay func for use where there is no timer tick

Code: [Select]
// Hang around for delay in ms. Approximate but doesn't need interrupts etc working.

static void hang_around(uint32_t delay)
{

uint32_t fred = 17000*delay;

while (fred>0)
{
fred--;
}

}

Putting volatile there makes a lot of stuff work :)

newbrain - the reason I post a lot is because I have nobody else I can ask. I am working more or less alone. I have nobody who is available as a "C consultant". I have actually tried to set something like that up (via my little business) but nobody is interested. I do a lot of googling of course but sometimes a focused Q on a good forum is a good way. I have always used forums, for various topics, not just electronics. But I am learning. It is a steep curve in places. I am documenting everything too, with lots of comments in the code (most progs don't comment C at all).
Yes, that's a typical case where an optimizer will completely remove the loop, and, as said, I'm enjoying this - so post away at yout leisure and need!
(not that I have any authority to say otherwise ;D)

Maybe there's one basic concept that needs to be mentioned:
C defines an "abstract machine" that does what you tell it to.
BUT, the only way one has to check that "abstract machine" is doing the right thing it's by its side effects outside of any particular function.
In this case fred is not visible or reachable outside of the function and its address is never taken, the loop does not modify any global object etc. etc.: the compiler is at freedom to rewrite the code as fred=0 (and then throw even that away!).
Note that the time it takes to do something or the fact that a debugger can probe the code are (in abstract) not of any concern to this "abstract machine".
Using the volatile type qualifier instruct the compiler to create code that (again, externally) completes all side effects before the volatile access, and completes the volatile access and its side effects before (again externally) performing other side effects.

I've mostly refrained to quote the standard, but this is how it's described in C11, I think it's quite clear (emphasis mine):
Quote
5.1.2.3  Program execution
1 The  semantic  descriptions  in  this  International  Standard  describe  the  behavior  of  an abstract machine in which issues of optimization are irrelevant.
2 Accessing  a  volatile  object,  modifying  an  object,  modifying  a  file,  or  calling  a  function that does any of those operations are all side effects, which are changes in the state of the  execution  environment. Evaluation of  an  expression  in  general  includes  both  value computations  and  initiation  of  side  effects. Value  computation  for  an  lvalue  expression includes determining the identity of the designated object.
3 [ I'll spare you the definition of sequencing ]
4 In  the  abstract  machine,  all  expressions  are  evaluated  as  specified  by  the  semantics. An actual  implementation  need  not  evaluate  part  of  an  expression  if  it  can  deduce  that  its value is not used and that no needed side effects are produced (including any caused by calling a function or accessing a volatile object).
5 [ signal handling ]
6 The least requirements on a conforming implementation are:
— Accesses to volatile objects are evaluated strictly according to the rules of the abstract machine.
— At program termination, all data written into files shall be identical to the result that execution of the program according to the abstract semantics would have produced.
— The  input  and  output  dynamics  of  interactive  devices  shall  take place  as  specified  in 7.21.3. The intent  of  these  requirements  is  that  unbuffered  or  line-buffered  output appear as soon as possible, to ensure that prompting messages actually appear prior to a program waiting for input.
This is the observable behavior of the program.

EtA: note that for embedded programs, the first clause in 6 applies, the other might or might not (usually not much).
Title: Re: GCC compiler optimisation
Post by: NorthGuy on August 12, 2021, 01:28:55 pm
... the vast majority of stuff online is garbage... all the way to code examples which could not have ever worked.

People copy things over. The call it "content" and they use various techniques to move them to the top of Google searches. Thus, complete garbage often finds way to the tops of Google searches. Also Google doesn't search what you ask anymore. Rather, they use AI to figure out what you really want to search for and show you that. Sometimes it's working, but if it doesn't you're doomed.

IMHO, it's better to use datasheets and specs. C standard may not be easy to read, but it'll give you general idea how things look from the perspective of the standard.
Title: Re: GCC compiler optimisation
Post by: bson on August 12, 2021, 07:50:06 pm
Is there a command-line switch to stop this interpretation of "standard" functions?  (globally, or on a per-function basis?)
Sure. "-fno-builtin" for all of them, and then there are flags like "-fno-builtin-memcpy".
Note that the arguments to memcpy() and friends are not declared volatile, so the calls may still be dropped if they're not needed.  (Or you get compile errors complaining about the loss of volatile.)  They may also have arguments declared "restrict", meaning the compiler will try to detect if they overlap.  For a function like this that would also be undesirable.

Better is to declare specific functions that embody the desired semantics since they differ from common usage, for example:
Code: [Select]
volatile void* vmemcpy(volatile void* dst, const volatile void* src, size_t len) {
   volatile char* dst0 = (volatile char*)dst;
   const volatile char* src0 = (const volatile char*)src;
   while (len-- > 0) { *dst0++ = *src0++; } // or asm volatile (...);
   return dst;
}
Instead of replacing memcpy(), which is only confusing since it already has well-understood behavior.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 12, 2021, 09:58:54 pm
I found this works exactly right

Code: [Select]
// Hang around for delay in ms. Approximate but doesn't need interrupts etc working.

__attribute__((noinline))
static void hang_around(uint32_t delay)
{
  delay *= (SystemCoreClock/4000);

  asm volatile (
    "1: subs %[delay], %[delay], #1 \n"
    "   nop \n"
    "   bne 1b \n"
    : [delay] "+l"(delay)
  );
}

for FLASH, and for RAM based code the /4000 has to be /6000. Yes; surprised me too, since I thought RAM based code runs at a genuine 0WS while FLASH code runs at 0WS if you believe the ST story about the ART :)

SystemCoreClock is set elsewhere with SystemCoreClock=168000000; because obviously the CPU can't have any way to determine its clock speed :) Well, it could use the camera interface to scan the text on the crystal and do OCR on it...

As regards other dodgy code, I wonder about this

Code: [Select]
uint32_t offset=0;
uint32_t addr=0x08000000;

for ( uint32_t page=4100; page<=5119; page++ )
{
AT45dbxx_WritePage((uint8_t*)addr,512,page*512);
offset+=512;
addr+=512;
}

where addr is reading the CPU FLASH. Addr is being modified in the loop so could this really be optimised out?

I then have this bit which tests the uppermost (128k) block of FLASH for erasure and programming. That is referencing FLASH memory addresses, but they are being declared volatile, but I wonder if correctly

Code: [Select]


/**
  * @brief  Program word (32-bit) at a specified address.
  * @note   This function must be used when the device voltage range is from
  *         2.7V to 3.6V.
  *
  * @note   If an erase and a program operations are requested simultaneously,
  *         the erase operation is performed before the program one.
  *
  * @param  Address specifies the address to be programmed.
  * @param  Data specifies the data to be programmed.
  * @retval None
  * Waits for previous operation to finish
  *
  */

static void L_FLASH_Program_Word(uint32_t Address, uint32_t Data)
{
// wait for any previous op to finish
while(__HAL_FLASH_GET_FLAG(FLASH_FLAG_BSY) != RESET);
// clear program size bits
CLEAR_BIT(FLASH->CR, FLASH_CR_PSIZE);
// reload program size bits
FLASH->CR |= FLASH_PSIZE_WORD;
// enable programming
FLASH->CR |= FLASH_CR_PG;
// write the data in
*(volatile uint32_t*)Address = Data;
}


// ===== Make sure we can erase and program sector 11 - the top one =====
// If that works, the rest should work :)

uint32_t error=0;
uint32_t data;
uint32_t address;

// Erase sector
L_HAL_FLASH_Unlock();
L_FLASH_Erase_Sector(11);
L_HAL_FLASH_Lock();

// Check it is all FFs
for (address=0x080e0000; address<0x080fffff; address+=4)
{
data=*(volatile uint32_t*)address;
if ( data != 0xffffffff )
error++;
}

// Erase sector again
L_HAL_FLASH_Unlock();
L_FLASH_Erase_Sector(11);

// Fill it with data
for (uint32_t i=0; i<(128*1024); i+=4)
{
L_FLASH_Program_Word(i+0x080e0000, i);
}

// Probably always best to lock the flash again before reading it
L_HAL_FLASH_Lock();

// Check the data we have just written
data=0;
for (address=0x080e0000; address<0x080fffff; address+=4)
{
if ( (*(volatile uint32_t*)address) != data )
error++;
data+=4;
}

On the other matter, which was variables being held in registers and thus the debug mode being unable to display their values as you step through the code, am I right that this is rather useless (because half the variables you wanted to watch are not visible anymore, with -Og) but the only way around it is to use -O0?

I will look into those mem functions. I replaced memcpy with memmove (to shuffle the 7 byte string 1 byte left i.e. src and dest overlap) and it appears to be running fine.

Title: Re: GCC compiler optimisation
Post by: ataradov on August 12, 2021, 10:10:05 pm
As regards other dodgy code, I wonder about this
.....
where addr is reading the CPU FLASH. Addr is being modified in the loop so could this be optimised out?

No, there is no way that could be optimized, assuming AT45dbxx_WritePage() ends in SPI access.


I then have this bit which tests the uppermost (128k) block of FLASH for erasure and programming. That is referencing FLASH memory addresses, but they are being declared volatile, but I wonder if correctly
The code looks fine to me.


On the other matter, which was variables being held in registers and thus the debug mode being unable to display their values as you step through the code, am I right that this is rather useless (because half the variables you wanted to watch are not visible anymore, with -Og) but the only way around it is to use -O0?
It depends on the debugger. I'm not sure if it is possible to make Eclipse show the value of variables held in registers. Low level debug info has this information.

But as a workaround when dealing with stuff like this, I just make dummy global variables and assign the values I need to see to those variables.

Title: Re: GCC compiler optimisation
Post by: newbrain on August 12, 2021, 11:25:23 pm
They may also have arguments declared "restrict", meaning the compiler will try to detect if they overlap.  For a function like this that would also be undesirable.
Quite the contrary.
The 'restrict' type qualifier is a contract that bounds the objects not to overlap*.

The compiler relies on the contract to be honoured, and this allows better optimizations.

If you break the contract, the program is broken (that is: Undefined Behaviour), there's no need for the compiler to check.

*In this simple case.
The complete definition is more complicated (C11: 6.7.3§8, 6.7.3.1).
Slightly more precisely, it says that an object pointed by a restrict pointer is not access through any other pointers (for the duration of the restrict pointer lifetime).
Title: Re: GCC compiler optimisation
Post by: peter-h on August 13, 2021, 10:27:53 am
Is there some way to see what the compiler is removing? I think this came up before and basically the answer was No.

I've come up against a much more complicated problem: the USB logical drive won't even format. Basically this is a removable block device implementation supplied by ST, which calls just two functions that interface onto the serial FLASH: write block (with transparent erase) and read block. The whole thing is interrupt driven, though it starts off as an RTOS thread. It works with -O0 but not with -Og. The two FLASH funcs are widely used elsewhere and appear to work fine, which leaves a large chunk of impenetrable USB code... The two funcs are these but they call other stuff, some of which is barely penetrable ST lib code

Code: [Select]
// Write a number of bytes to a single page through the buffer with built-in erase.
// 1/8/21 The last parm is actually a linear address within the device.
bool AT45dbxx_WritePage(
const uint8_t *data, // In Data to write
uint16_t len, // In Length of data to write (in bytes)
uint32_t page // In linear address, a multiple of 512
) {

HAL_StatusTypeDef status = HAL_OK;

if (len==0) return (status == HAL_OK);

page = page << AT45dbxx.Shift;
at45dbxx_resume();
at45dbxx_wait_busy();
B_HAL_GPIO_WritePin(_45DBXX_CS_GPIO, _45DBXX_CS_PIN, GPIO_PIN_RESET);
at45dbxx_tx_rx_byte(AT45DB_MNTHRUBF1);
at45dbxx_tx_rx_byte((page >> 16) & 0xff);
at45dbxx_tx_rx_byte((page >> 8) & 0xff);
at45dbxx_tx_rx_byte(page & 0xff);
status = B_HAL_SPI_Transmit(&_45DBXX_SPI, (uint8_t *) data, len, AT45DB_SPI_TIMEOUT_MS);
B_HAL_GPIO_WritePin(_45DBXX_CS_GPIO, _45DBXX_CS_PIN, GPIO_PIN_SET);
at45dbxx_wait_busy();

return (status == HAL_OK);
}


// Read a number of bytes from a single page
// 1/8/21 The last parm is actually a linear address within the device. May have to be a
// multiple of 512, too.
// The device command used (0Bh) allows a continuous read of the whole device. This does not
// appear to be allowed here.

bool AT45dbxx_ReadPage(
uint8_t *data, // Out Buffer to read data to
uint16_t len, // In Length of data to read (in bytes)
uint32_t page // In linear address, a multiple of 512
) {
HAL_StatusTypeDef status = HAL_OK;

if (len==0) return (status == HAL_OK);

page = page << AT45dbxx.Shift;

// Round down length to the page size
if (len > AT45dbxx.PageSize) {
len = AT45dbxx.PageSize;
}

at45dbxx_resume();
at45dbxx_wait_busy();
B_HAL_GPIO_WritePin(_45DBXX_CS_GPIO, _45DBXX_CS_PIN, GPIO_PIN_RESET);
at45dbxx_tx_rx_byte(AT45DB_RDARRAYHF);
at45dbxx_tx_rx_byte((page >> 16) & 0xff);
at45dbxx_tx_rx_byte((page >> 8) & 0xff);
at45dbxx_tx_rx_byte(page & 0xff);
at45dbxx_tx_rx_byte(0);
status = B_HAL_SPI_Receive(&_45DBXX_SPI, data, len, AT45DB_SPI_TIMEOUT_MS);
B_HAL_GPIO_WritePin(_45DBXX_CS_GPIO, _45DBXX_CS_PIN, GPIO_PIN_SET);

return (status == HAL_OK);
}

I can be very sure that all my "IO" references are volatile because I only ever use the ST defs and all are prefixed with __IO which is #defined as volatile.

To be honest I am happy to just continue the project with -O0 because it runs easily fast enough. It still removes code which manifestly never gets reached e.g. anything after a for( ;; ); but that's ok; I should not actually have any of that (except when testing when you want to insert say an LED flash loop near the start of main.c but don't want the other 99% of the product, including the RTOS and all the timer ISRs, stripped out ;) and then you have to fool the compiler by decrementing a unit32_t around that loop). With the other optimisation options I was seeing weird stuff where e.g. I had a variety of if() tests in a sequence, and at some point it decided that the expression being tested had to be always false, but I could not see that at all. Obviously the compiler could be way more clever than me in detecting an "impossible-true" situation, but the fact is that all those conditionals were working correctly. I think the compiler was deciding that some flags (various bytes in a 512-byte buffer read from an eeprom) would always test false. Maybe that whole buffer should have been volatile, but that would not make sense to me if it fixed it. The bigger issue is that so much code could be broken in this way and retesting absolutely everything one has written in 6 months is just not viable. And some 90% of this project is libs from ST and such (ETH and USB) which are practically impenetrable and take months to do any debugging on, and debugging -O issues is very slow since it involves stepping through code, which is often impossible because the code is real-time.

The code size with -O0 and -Og is 220k and 160k which is quite a lot and suggests that stuff is being removed whole, somewhere.

Maybe there is a specific compiler option one could try to narrow it down? Where would these be specified in Cube?

As an aside, could optimisation break code where you call a function which has say 4 parameters but the 4th one is never used inside that function?
Title: Re: GCC compiler optimisation
Post by: Siwastaja on August 13, 2021, 12:15:57 pm
My suggestion: always compile with -O2, -O3 or -Os. If your program does not work, don't try if it would with -O0. Also don't try to look what code is "removed" by optimizer. Looking at assembly listing is sometimes needed, but shouldn't be your default strategy, either.

The reason why the code does not work is that the code is wrong. It's buggy. You need to apply all the "normal" debugging strategies to find the reason. Adding logging or facilitating debugger features brings you a long way.

You seem to be fixated to the idea that your program is fine or at least "almost" fine and it's the compiler that breaks it and then you need to kind of reverse-engineer what the compiler is doing to "break" the program then "adjust" the program until it gets "through" the compiler. This is all wrong, don't do it. But your way of debugging the issues amplifies this false premise.

Try more usual ways of dealing with bugs, even if they are bugs that only appear at certain compiler settings. If you decide not to mess up with the settings all the time for no reason, you simply don't know, and in this case it'd be a bliss, allowing you to treat bugs as bugs.

In other words, optimization is a tiny detail in how compiler works, but broken code is broken regardless of compiler settings - even if some setting, like -O0, by accident seemingly fixes the problem. By focusing on the optimization settings that are not the issue at all, you just waste your time.
Title: Re: GCC compiler optimisation
Post by: newbrain on August 13, 2021, 12:29:55 pm
As an aside, could optimisation break code where you call a function which has say 4 parameters but the 4th one is never used inside that function?
Remember that optimization only breaks broken code - but having ST lib in the recipe, that is quite possible.
You might try to not optimize only those parts by changing the C/C++ properties on the relevant folders (or files) in the project view, IIRC (yes that works, I checked).

The difference in size going from -O0 to -Og is similar to what I get on a project of similar volume (different MCU, an NXP iMX RT 1021), so no surprise there.

The read and write page functions look OK - when making some assumption on the parts not shown.

You seem to be fixated to the idea that your program is fine or at least "almost" fine and it's the compiler that breaks it and then you need to kind of reverse-engineer what the compiler is doing to "break" the program then "adjust" the program until it gets "through" the compiler. This is all wrong, don't do it. But your way of debugging the issues amplifies this false premise.
Yes, very much this. It's a backasswards strategy.

I personally have yet to encounter a real compiler bug - of course they exist but the usual suspects (gcc, clang and even cl) are quite good.
Title: Re: GCC compiler optimisation
Post by: NorthGuy on August 13, 2021, 01:23:06 pm
Is there some way to see what the compiler is removing? I think this came up before and basically the answer was No.

Sure. Look at the assembler generated (--save-temps switch in GCC). It'll show you exactly what compiler did.

However, this is not a good idea. You cannot rely on this. The code generated may change any time. Write to the standard and let the compiler do code generation.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 13, 2021, 02:44:48 pm
OK, yes, if I was writing something from scratch, then why not use O2 or O3. Then if something doesn't run you go straight to the last bit you did.

But here I have a load of code, most of it 3rd party libs, no support on any of it of course...

"Remember that optimization only breaks broken code - but having ST lib in the recipe, that is quite possible."

Exactly...
Title: Re: GCC compiler optimisation
Post by: ejeffrey on August 13, 2021, 03:44:11 pm
One more thing to look for: you say ISRs get removed if you add an infinite loop to main.  If you are talking about ISR functions installed through a runtime function that makes sense.  If you are talking about the ISRs in the default flash interrupt table, that shouldn't happen.  The default ISR table is supposed to be marked as "keep" in the linker script so that it serves as a root for the --gc-sections option.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 13, 2021, 03:53:02 pm
What I meant is that the remainder of main.c would be stripped out.

But perhaps this is wrong. The compiler should strip out only the remainder of the function main().

However, that is not what happened. My program lost about 90% of its size :)
Title: Re: GCC compiler optimisation
Post by: ataradov on August 13, 2021, 04:29:41 pm
I agree with everyone else, forget that -O0 exists. It is never useful in real life. This way when something "breaks" you will have a limited scope of changes since the last working build to review.

If you are already facing the project that is broken like this, then you need to start looking at disassembly and check what was optimized.
Title: Re: GCC compiler optimisation
Post by: Siwastaja on August 13, 2021, 04:46:19 pm
Debugging code made by others, including external "libraries" (I'm assuming you don't get object files + headers but actual source code because you talk about compiling it), is no different at all from debugging your own code. Use the same strategies.

If compiler completely removes some code, this is only because the code does nothing. Look at the code, figure out what it is supposed to do, then fix it so that it does that.

Yes this is tedious work, writing code is, but this is the only way, you can't fix it by trying different compiler settings, looking at disassembly and post about the stupidness of all this on the forum.

Sometimes it's faster to write completely from scratch than to decipher and fix non-working codebases that are basically untested.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 13, 2021, 07:10:35 pm

I have been very carefully working through it. The bit which isn't working has been narrowed down but is still a lot of code, from ST mainly, and with lots of pointers and such. How do I select -O0 on a particular .c file only? Is there some attribute one can put in the start of the file?
Title: Re: GCC compiler optimisation
Post by: ejeffrey on August 13, 2021, 08:03:51 pm
What I meant is that the remainder of main.c would be stripped out.

But perhaps this is wrong. The compiler should strip out only the remainder of the function main().

However, that is not what happened. My program lost about 90% of its size :)

If you use -ffunction-sections and -fgc-sections then each function will be placed in its own section, and the linker (not compiler!) will do a garbage collection pass to remove unreferenced functions.  If you do this (and it is very standard on microcontroller builds) then adding an infinite loop in main will cause everything except the init code to be removed.  Which is not a problem at all since that code will never be called.  Looking at what code is removed is counter productive.  You should only look at where your code *behaves* improperly.  Fix that, and either the rest of your code will "come back", or it isn't needed.  Fixating on the "missing" code is 99% of the time a red herring.
Title: Re: GCC compiler optimisation
Post by: langwadt on August 13, 2021, 08:14:37 pm
if code written by hammering the keyboard and "fixing it" with chewing gum and duct tape doesn't work it is obviously the libraries or the compilers fault ;)
Title: Re: GCC compiler optimisation
Post by: peter-h on August 13, 2021, 08:49:38 pm
I narrowed it down to a couple of files, and setting both to -O0 makes the whole thing run ok. One of them is FatFS and the other is a load of mainly ST code for serial FLASH and SPI. I don't think the root issue is in FatFS - partly because lots of people use it, and partly because the problem also shows up when trying to format the block device from Windows and that doesn't use FatFS (which merely enables internal code to see the filesystem); it uses just the latter stuff (ST mostly).

Unfortunately, in ST Cube at least, -Og makes it very hard to do debugging, because literally half the variables cannot be viewed - because they have been moved to registers. One can switch to register mode and by reference to the disassembly listing it is usually possible to see the value of the variable, but it's quite clumsy. I wonder if this is debugger dependent? I am using STLINK V2 & V3, and have a Segger Edu kicking around somewhere. Unfortunately there are also many cases where I pass the address of a buffer to a function, and cannot view its contents because it also says "optimised out".

On the plus side, it looks like all the code I have written myself is working fine with -Og :)

In Cube, one can select an -O level on a per-file basis, simply by right-clicking on that file and going to Properties. A tiny symbol appears next to the file, which disappears if you build the whole project with the same -O level as that file (so you have to be fairly careful there).

It's been a great learning experience - thank you all :)
Title: Re: GCC compiler optimisation
Post by: westfw on August 13, 2021, 10:03:40 pm
Quote
I have a load of code, most of it 3rd party libs, no support on any of it of course...   :
 The bit which isn't working has been narrowed down but is still a lot of code, from ST mainly
Surely code from ST has SOME support?

Quote
demonstrate GCC optimising out a straight read of 0x08000000
Yes: https://godbolt.org/z/he6avac4E (https://godbolt.org/z/he6avac4E)
That example has:

Code: [Select]
void f()
{
    char buffer[512];
    memcpy(buffer, (char*)0x08000000, 512);
}
If I change that code to have "volatile char buffer[512];", the copy should no longer be optimized away, right?
But it looks like it is (using gcc 11 and higher.  gcc 10 creates actual code...)  ARM gcc behaves similarly...

Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 13, 2021, 10:18:05 pm
Code: [Select]
void f()
{
    char buffer[512];
    memcpy(buffer, (char*)0x08000000, 512);
}
If I change that code to have "volatile char buffer[512];", the copy should no longer be optimized away, right?
But it looks like it is (using gcc 11 and higher.  gcc 10 creates actual code...)  ARM gcc behaves similarly...

Well, no.
As we talked about way earlier in the thread, the memcpy() function itself doesn't have volatile-qualified parameters.
So when you're passing a volatile [] to memcpy(), the parameter is converted to a pointer to non-volatile. So it doesn't make a difference. That's why I said earlier that the only way of solving this would be to write your own memcpy() function.

If you assign stuff to the local buffer array though within this function, it will make the difference you expect.
As in:
Code: [Select]
void f()
{
    volatile char buffer[512];
    memcpy(buffer, (char*)0x08000000, 512); // still pruned
    buffer[0] = 1; // not pruned
}

Interestingly, if you swap the destination and the source, then the memcpy() doesn't get pruned, even with no volatile qualifier:
Code: [Select]
void f()
{
    char buffer[512];
    memcpy((char*)0x08000000, buffer, 512);
}

The compiler probably allows itself to make more assumptions for objects that it knows about than for objects that it doesn't (here, an absolute address). Playing a bit with this will exhibit even funnier stuff...

Of course don't take all this as rules - this is just implementation-dependent behavior. Use volatile as needed.
Title: Re: GCC compiler optimisation
Post by: westfw on August 13, 2021, 11:40:15 pm
BTW, I've completely lost track of what we're talking about.
The OP seemed to be about bootloaders and whether bootloaded code should contain startup code (yes, it should!  My philosophy is that code that is bootloaded should look exactly like code that is used without a bootloader, except for its position.  It should still have the initial SP and PC and the rest of the vectors, and still have all of the stuff that happens before main() is called.  (You're not still having the bootloader go directly to your application main() without doing the initialized variable copy and stuff, are you?  That could explain some issues!))

Then we diverted to a discussion of compiler optimization and behavior of "volatile."
There is code that is "90% smaller with optimization" and code which "doesn't work", but it's not clear that they're the same code...


Quote
-Og makes it very hard to do debugging, because literally half the variables cannot be viewed - because they have been moved to registers.
If you're looking for code that has been completely omitted, you'd be debugging code flow rather than variable content, wouldn't you?
If what you think are global ram variables cannot be viewed, that's a pretty substantial clue right there...


Quote
wanted 500/month for ongoing "support" which is way too much for what will be needed (maybe 1hr/month).

That actually doesn't seem all that unreasonable.  1hr of actual work in a month probably means another hour or two worth of re-familiarizing yourself with code and "interfacing with customer", and per-hour charges in the $300 range aren't uncommon (although, you said Pounds, I think...)
https://www.fullstacklabs.co/blog/software-development-price-guide-hourly-rate-comparison (https://www.fullstacklabs.co/blog/software-development-price-guide-hourly-rate-comparison)

Title: Re: GCC compiler optimisation
Post by: westfw on August 13, 2021, 11:58:33 pm
Quote
As we talked about way earlier in the thread, the memcpy() function itself doesn't have volatile-qualified parameters.
Hmm.  But it's only because the compiler has internal knowledge of what memcpy() is supposed to do.  Normally it would just be an external function with unknown side-effects, and it would be called with the given arguments.  If the compiler is going to treat it like a "language feature" instead of a library function, it should take into account the additional semantics that it is (should be) aware of.  (and it did so up until gcc 11, apparently, even when it used inline code for the copy...)
(I guess this is a fine example of "I don't like it, so I think it's wrong.")
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 14, 2021, 12:19:22 am
Quote
As we talked about way earlier in the thread, the memcpy() function itself doesn't have volatile-qualified parameters.
Hmm.  But it's only because the compiler has internal knowledge of what memcpy() is supposed to do.

Of course, technically this is why. This knowledge allows compilers to generate pretty efficient inline code for those functions, tailored to the task at hand.

Normally it would just be an external function with unknown side-effects, and it would be called with the given arguments.

Well, if the compiler has no knowledge of the possible side-effects of a given function, of course it must call it. As it happens though, modern compilers have internal knowledge of the most common (if not all? I dont know for sure) functions from the C std library, which allows them to implement more efficient code. There are tons of examples. One with the printf() function: called with a string without format specifiers (when the compiler can statically see this), GCC (and probably CLANG) will just call puts() instead.

I admit this can be confusing to many. The C std library becomes an integral part of the compiler. While this seems reasonable, this can lead to misconceptions and actual issues, especially on systems for which the C std lib is implemented in dynamic libraries, in which case, for a given C std lib function call, you may end up either with an inline, compiler-dependent version, or with a call to an export in a dynamic library, possibly implementing the same function in a slightly different way...
Title: Re: GCC compiler optimisation
Post by: bson on August 14, 2021, 12:59:35 am
That example has:

Code: [Select]
void f()
{
    char buffer[512];
    memcpy(buffer, (char*)0x08000000, 512);
}
If I change that code to have "volatile char buffer[512];", the copy should no longer be optimized away, right?
But it looks like it is (using gcc 11 and higher.  gcc 10 creates actual code...)  ARM gcc behaves similarly...
Probably some confusion over auto variables vs volatile.  It knows buffer is never used, so it and any code that is used to compute or initialize it can be removed.  The fact that it's on the stack means the compiler sees its entire lifespan, from function entry to exit, and knows it's never used.  Volatile should override that and make it non-removable, but it's such an odd thing to have local volatile variables that it's not surprising if there's debate over whether it can be optimized out of existence or not; it's really more about literal implementation of the language spec.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 14, 2021, 06:25:21 am
I think there are multiple things going on in my project, and probably in a couple of files. I did some very careful debugging. Compiled it all with -O0, formatted the disk, and placed a file on it, and verified it. It has a CRC on the end. Then from the inside (via FatFS) read it to check the CRC. All good. For those who like to see code, this is the CRC func

Code: [Select]


/*
*
* CRC-32/JAMCRC. This is a "rolling" algorithm. Invert the *final* result for ISO-HDLC.
* Returns 0x098494f3 from "123456789" (9 bytes). Inverting this (at the end) gives 0xcbf43926.
* Polynomial is 0x04c11db7 and it holds it backwards (0xEDB88320).
* This version accepts one byte at a time, and maintains the CRC in crcvalue. This makes it suitable
* for calculating a CRC across a number of data blocks.
* Speed is approx 400kbytes/sec.
* crcvalue must be initialised to 0xffffffff by caller, and holds the accumulated CRC as you go along
* See e.g. [url]https://crccalc.com/[/url] [url]https://www.lammertbies.nl/comm/info/crc-calculation[/url]
* [url]https://reveng.sourceforge.io/crc-catalogue/17plus.htm#crc.cat-bits.32[/url]
*
*/


void crc32(uint8_t input_byte, uint32_t *crcvalue)
{
      for (uint32_t j=0; j<8; j++)
      {
      uint32_t mask = (input_byte^*crcvalue) & 1;
      *crcvalue>>=1;
      if(mask)
      *crcvalue=*crcvalue^0xEDB88320;
      input_byte>>=1;
      }
}



/*
 *
 * Check CRC, stored in last 4 bytes of a file.
 * For speed, reads into a 512 byte buffer.
 * This code is tricky, due to CRC potentially split across buffer boundaries.
 * crcinit is normally initialised to 0xffffffff by the caller.
 * No upper limit on file size. Minimum size = 5 bytes.
 * Returns true if CRC is good, false otherwise (including if file doesn't exist)
 * Exec time for 1MB: 15s.
 *
 */

bool filecrc(char * filename, uint32_t crcinit)
{

FILINFO fno;

uint32_t crc=crcinit;
uint32_t filecrc=0;
uint32_t offset=0;
int32_t bytesleft;
uint8_t pagebuf[512];
uint32_t numread=0;

if ( KDE_get_file_properties ( filename, &fno ) == false ) return (false);

bytesleft=fno.fsize;

if ( bytesleft<5 ) return (false);

do
{
// Read up to 512 bytes into pagebuf
if ( KDE_file_read( filename, offset, 512, pagebuf, &numread) == false )
return (false);
offset+=512;
bytesleft-=numread;

if ((numread==512) && (bytesleft>=4))
// Most common case: 512 bytes read, not the last page, and no CRC bytes yet
{
for (int i=0; i<512; i++)
{
crc32(pagebuf[i], &crc);
}
}
else
{
if ((numread<=512) && (bytesleft==0))
// Last page, and contains entire CRC
{
for (int i=0; i<(numread-4); i++)
{
crc32(pagebuf[i], &crc);
}
filecrc=pagebuf[numread-4]|(pagebuf[numread-3]<<8)|(pagebuf[numread-2]<<16)|(pagebuf[numread-1]<<24);
}
else
{
// CRC is split (buffer contains only some of it). Calc CRC over all data in buffer
for (int i=0; i<(512+bytesleft-4); i++)
{
crc32(pagebuf[i], &crc);
}
// read just CRC bytes
if ( KDE_file_read( filename, offset+bytesleft-4, 512, pagebuf, &numread) == false )
return (false);
filecrc=pagebuf[0]|(pagebuf[1]<<8)|(pagebuf[2]<<16)|(pagebuf[3]<<24);
bytesleft=0; // ensure end of loop
}
}
}
while ( bytesleft > 0 );

crc=~crc;  // invert JAMCRC to HDLC CRC which the file has on the end
if ( crc==filecrc ) return (true);

return (false);
}

Then I compiled the whole thing with -Og. The CRC calculation fails. There is nothing obviously wrong with the data, and ideally I could spend time repeating this with a small file, say 30 bytes, to narrow it down (the test file is 1MB). So definitely one problem here.

But even before the CRC fails, if FatFS is also compiled with -Og, FatFS cannot find the file on the disk! So I stepped through the code, and narrowed it down to somewhere deep in FatFS returning false - around here

Code: [Select]

/*-----------------------------------------------------------------------*/
/* Directory handling - Find an object in the directory                  */
/*-----------------------------------------------------------------------*/

static
FRESULT dir_find ( /* FR_OK(0):succeeded, !=0:error */
DIR* dp /* Pointer to the directory object with the file name */
)
{
FRESULT res;
FATFS *fs = dp->obj.fs;
BYTE c;
#if _USE_LFN != 0
BYTE a, ord, sum;
#endif

res = dir_sdi(dp, 0); /* Rewind directory object */
if (res != FR_OK) return res;
#if _FS_EXFAT
if (fs->fs_type == FS_EXFAT) { /* On the exFAT volume */
BYTE nc;
UINT di, ni;
WORD hash = xname_sum(fs->lfnbuf); /* Hash value of the name to find */

while ((res = dir_read(dp, 0)) == FR_OK) { /* Read an item */
#if _MAX_LFN < 255
if (fs->dirbuf[XDIR_NumName] > _MAX_LFN) continue; /* Skip comparison if inaccessible object name */
#endif
if (ld_word(fs->dirbuf + XDIR_NameHash) != hash) continue; /* Skip comparison if hash mismatched */
for (nc = fs->dirbuf[XDIR_NumName], di = SZDIRE * 2, ni = 0; nc; nc--, di += 2, ni++) { /* Compare the name */
if ((di % SZDIRE) == 0) di += 2;
if (ff_wtoupper(ld_word(fs->dirbuf + di)) != ff_wtoupper(fs->lfnbuf[ni])) break;
}
if (nc == 0 && !fs->lfnbuf[ni]) break; /* Name matched? */
}
return res;
}
#endif
/* On the FAT12/16/32 volume */
#if _USE_LFN != 0
ord = sum = 0xFF; dp->blk_ofs = 0xFFFFFFFF; /* Reset LFN sequence */
#endif
do {
res = move_window(fs, dp->sect);
if (res != FR_OK) break;
c = dp->dir[DIR_Name];
if (c == 0) { res = FR_NO_FILE; break; } /* Reached to end of table */
#if _USE_LFN != 0 /* LFN configuration */
dp->obj.attr = a = dp->dir[DIR_Attr] & AM_MASK;
if (c == DDEM || ((a & AM_VOL) && a != AM_LFN)) { /* An entry without valid data */
ord = 0xFF; dp->blk_ofs = 0xFFFFFFFF; /* Reset LFN sequence */
} else {
if (a == AM_LFN) { /* An LFN entry is found */
if (!(dp->fn[NSFLAG] & NS_NOLFN)) {
if (c & LLEF) { /* Is it start of LFN sequence? */
sum = dp->dir[LDIR_Chksum];
c &= (BYTE)~LLEF; ord = c; /* LFN start order */
dp->blk_ofs = dp->dptr; /* Start offset of LFN */
}
/* Check validity of the LFN entry and compare it with given name */
ord = (c == ord && sum == dp->dir[LDIR_Chksum] && cmp_lfn(fs->lfnbuf, dp->dir)) ? ord - 1 : 0xFF;
}
} else { /* An SFN entry is found */
if (!ord && sum == sum_sfn(dp->dir)) break; /* LFN matched? */
if (!(dp->fn[NSFLAG] & NS_LOSS) && !mem_cmp(dp->dir, dp->fn, 11)) break; /* SFN matched? */
ord = 0xFF; dp->blk_ofs = 0xFFFFFFFF; /* Reset LFN sequence */
}
}
#else /* Non LFN configuration */
dp->obj.attr = dp->dir[DIR_Attr] & AM_MASK;
if (!(dp->dir[DIR_Attr] & AM_VOL) && !mem_cmp(dp->dir, dp->fn, 11)) break; /* Is it a valid entry? */
#endif
res = dir_next(dp, 0); /* Next entry */
} while (res == FR_OK);

return res;
}

Unfortunately with -Og half the variables are "optimised out" so difficult to work out what is what.

The "unable to format" from Windows is a different issue however. This is nothing to do with FatFS. It is probably low down in the ST stuff like this

Code: [Select]

HAL_StatusTypeDef B_HAL_SPI_Transmit(SPI_HandleTypeDef *hspi, uint8_t *pData, uint16_t Size, uint32_t Timeout)
{
//  uint32_t tickstart;
  HAL_StatusTypeDef errorcode = HAL_OK;
  uint16_t initial_TxXferCount;

  initial_TxXferCount = Size;

  if (hspi->State != HAL_SPI_STATE_READY)
  {
    errorcode = HAL_BUSY;
    goto error;
  }

  if ((pData == NULL) || (Size == 0U))
  {
    errorcode = HAL_ERROR;
    goto error;
  }

  /* Set the transaction information */
  hspi->State       = HAL_SPI_STATE_BUSY_TX;
  hspi->ErrorCode   = HAL_SPI_ERROR_NONE;
  hspi->pTxBuffPtr  = (uint8_t *)pData;
  hspi->TxXferSize  = Size;
  hspi->TxXferCount = Size;

  /*Init field not used in handle to zero */
  hspi->pRxBuffPtr  = (uint8_t *)NULL;
  hspi->RxXferSize  = 0U;
  hspi->RxXferCount = 0U;
  hspi->TxISR       = NULL;
  hspi->RxISR       = NULL;

  /* Configure communication direction : 1Line */
  if (hspi->Init.Direction == SPI_DIRECTION_1LINE)
  {
    SPI_1LINE_TX(hspi);
  }

  /* Check if the SPI is already enabled */
  if ((hspi->Instance->CR1 & SPI_CR1_SPE) != SPI_CR1_SPE)
  {
    /* Enable SPI peripheral */
    __HAL_SPI_ENABLE(hspi);
  }

  /* Transmit data in 8 Bit mode */

    if ((hspi->Init.Mode == SPI_MODE_SLAVE) || (initial_TxXferCount == 0x01U))
    {
      *((__IO uint8_t *)&hspi->Instance->DR) = (*hspi->pTxBuffPtr);
      hspi->pTxBuffPtr += sizeof(uint8_t);
      hspi->TxXferCount--;
    }
    while (hspi->TxXferCount > 0U)
    {
      /* Wait until TXE flag is set to send data */
      if (__HAL_SPI_GET_FLAG(hspi, SPI_FLAG_TXE))
      {
        *((__IO uint8_t *)&hspi->Instance->DR) = (*hspi->pTxBuffPtr);
        hspi->pTxBuffPtr += sizeof(uint8_t);
        hspi->TxXferCount--;
      }
    }


  /* Clear overrun flag in 2 Lines communication mode because received is not read */
  if (hspi->Init.Direction == SPI_DIRECTION_2LINES)
  {
    __HAL_SPI_CLEAR_OVRFLAG(hspi);
  }

  if (hspi->ErrorCode != HAL_SPI_ERROR_NONE)
  {
    errorcode = HAL_ERROR;
  }

error:
  hspi->State = HAL_SPI_STATE_READY;
  return errorcode;
}


but there is something interesting in there which may be a clue, but which is above my pay grade to understand the code. It is this

Code: [Select]
       *((__IO uint8_t *)&hspi->Instance->DR) = (*hspi->pTxBuffPtr);
__IO is #defined as volatile, but DR itself also is in the ST .h files. They appear to be creating a pointer to the hspi structure, whose member DR is the SPI data register. Yet the vast majority of references to hspi is not "volatile" qualified, and I don't know why it should be. OTOH the compiler sees hspi whole and may decide to prune unreferenced members (of which there are plenty; ST love a structure for absolutely everything). It could then re-pack it, but that should still work, eh?

Funnily enough FatFS has its own private memcpy. They call it mem_cpy. And others. Maybe they know something:

Code: [Select]
/*-----------------------------------------------------------------------*/
/* String functions                                                      */
/*-----------------------------------------------------------------------*/

/* Copy memory to memory */
static
void mem_cpy (void* dst, const void* src, UINT cnt) {
BYTE *d = (BYTE*)dst;
const BYTE *s = (const BYTE*)src;

if (cnt) {
do {
*d++ = *s++;
} while (--cnt);
}
}

/* Fill memory block */
static
void mem_set (void* dst, int val, UINT cnt) {
BYTE *d = (BYTE*)dst;

do {
*d++ = (BYTE)val;
} while (--cnt);
}

/* Compare memory block */
static
int mem_cmp (const void* dst, const void* src, UINT cnt) { /* ZR:same, NZ:different */
const BYTE *d = (const BYTE *)dst, *s = (const BYTE *)src;
int r = 0;

do {
r = *d++ - *s++;
} while (--cnt && r == 0);

return r;
}

/* Check if chr is contained in the string */
static
int chk_chr (const char* str, int chr) { /* NZ:contained, ZR:not contained */
while (*str && *str != chr) str++;
return *str;
}



Title: Re: GCC compiler optimisation
Post by: gf on August 14, 2021, 07:07:16 am
That example has:

Code: [Select]
void f()
{
    char buffer[512];
    memcpy(buffer, (char*)0x08000000, 512);
}
If I change that code to have "volatile char buffer[512];", the copy should no longer be optimized away, right?
But it looks like it is (using gcc 11 and higher.  gcc 10 creates actual code...)  ARM gcc behaves similarly...
Probably some confusion over auto variables vs volatile.  It knows buffer is never used, so it and any code that is used to compute or initialize it can be removed.  The fact that it's on the stack means the compiler sees its entire lifespan, from function entry to exit, and knows it's never used.  Volatile should override that and make it non-removable, but it's such an odd thing to have local volatile variables that it's not surprising if there's debate over whether it can be optimized out of existence or not; it's really more about literal implementation of the language spec.

The volatile qualifier has virtually no meaning for the variable itself.
It only has an effect when a memory location (i.e. either the address of a variable, or a pointer) is dereferenced in the code flow.
If the dereferenced address (pointer) refers to a volatile memory location, then the load/store to the memory location cannot be optimized out.

Examples:

Code: [Select]
volatile int a;
a = 5;         // <== volatile memory store, because &a is a pointer to a volatile memory location

Code: [Select]
int a;
volatile int *__tmp = &a;
*__tmp = 5;    // <== volatile memory store, because __tmp is a pointer to a volatile memory location

Code: [Select]
volatile int a;
int *p = (int*)&a;
*p = 5;        // <== NOT a volatile memory store, alhtough it assigns 5 to a, and a was declared volatile

The effect of volatile is only that the instructions for a = 5 and *__tmp = 5 cannot be eliminated by the optimizer 1).
Retaining the variable a is just the consequence of not eliminating a = 5 in the first place, because a is used, then.
But if a were not used, it still can be eliminated, regardless whether qualified volatile or not.


1) In addition, the standard also prohibits re-ordering - for details please refer directly to the standard.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 14, 2021, 07:12:15 am
Don't look at FatFS, it is used by so many people that it is guaranteed to work with any level of optimizations. Look closer at the difference in behaviour of low level functions that actually read and write sectors.
Title: Re: GCC compiler optimisation
Post by: westfw on August 14, 2021, 08:10:56 am
Quote
The fact that it's on the stack means the compiler sees its entire lifespan, from function entry to exit, and knows it's never used.
Hmm.  An interesting theory.
If I change the buffer to "static volatile", then gcc will produce code to do the copy.  But clang doesn't. :-)
I like the "memcpy() doesn't know about volatile" explanation better, technically.

It's nice that I can turn off the "builtin" functions; can I do that on a per-call basis?
I guess I can use "__builtin_memcpy()" if I want to use the "optimized, built-in" version and have builtins otherwise turned off.  What about the other way around?(Hmm.  Clang seems to have a function attribute __attribute__((no_builtin("memcpy"))) that is supposed to last for the scope of the function with the attribute, but ... it doesn't seem to work on the example we're using :-( )  It must be deciding that buffer is useless even if it IS volatile and static.
https://clang.llvm.org/docs/AttributeReference.html#no-builtin
Title: Re: GCC compiler optimisation
Post by: peter-h on August 14, 2021, 08:17:17 am
"Look closer at the difference in behaviour of low level functions that actually read and write sectors."

Indeed, and the disk format failure is a clue (FatFS is not used). I posted the ST code  for the SPI stuff above. But debugging the optimised version is just too hard.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 14, 2021, 08:47:38 am
I posted the ST code  for the SPI stuff above.
That is just the transmit part. And it is way too low level. You need to look at the whole sector read, including setting of the address.

And then you need to look at how your chip select is formed for SPI. Your optimized code may be too fast and it breaks setup time for the CS before data transfer. And your SPI code may be too fast that your memory can't handle it.

Find the call that FatFS does to read a sector and start from that, not just the lowest level code.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 14, 2021, 09:28:36 am
These are the sector I/O

Code: [Select]
// Write a number of bytes to a single page through the buffer with built-in erase.

bool AT45dbxx_WritePage(
const uint8_t *data, // In Data to write
uint16_t len, // In Length of data to write (in bytes)
uint32_t page // In linear address, a multiple of 512
) {

HAL_StatusTypeDef status = HAL_OK;

if (len==0) return (status == HAL_OK);

page = page << AT45dbxx.Shift;
at45dbxx_resume();
at45dbxx_wait_busy();
B_HAL_GPIO_WritePin(_45DBXX_CS_GPIO, _45DBXX_CS_PIN, GPIO_PIN_RESET);
at45dbxx_tx_rx_byte(AT45DB_MNTHRUBF1);
at45dbxx_tx_rx_byte((page >> 16) & 0xff);
at45dbxx_tx_rx_byte((page >> 8) & 0xff);
at45dbxx_tx_rx_byte(page & 0xff);
status = B_HAL_SPI_Transmit(&_45DBXX_SPI, (uint8_t *) data, len);
B_HAL_GPIO_WritePin(_45DBXX_CS_GPIO, _45DBXX_CS_PIN, GPIO_PIN_SET);
at45dbxx_wait_busy();

return (status == HAL_OK);
}


// Read a number of bytes from a single page
// 1/8/21 The last parm is actually a linear address within the device.

bool AT45dbxx_ReadPage(
uint8_t *data, // Out Buffer to read data to
uint16_t len, // In Length of data to read (in bytes)
uint32_t page // In linear address, a multiple of 512
) {
HAL_StatusTypeDef status = HAL_OK;

if (len==0) return (status == HAL_OK);

page = page << AT45dbxx.Shift;

// Round down length to the page size
if (len > AT45dbxx.PageSize) {
len = AT45dbxx.PageSize;
}

at45dbxx_resume();
at45dbxx_wait_busy();
B_HAL_GPIO_WritePin(_45DBXX_CS_GPIO, _45DBXX_CS_PIN, GPIO_PIN_RESET);
at45dbxx_tx_rx_byte(AT45DB_RDARRAYHF);
at45dbxx_tx_rx_byte((page >> 16) & 0xff);
at45dbxx_tx_rx_byte((page >> 8) & 0xff);
at45dbxx_tx_rx_byte(page & 0xff);
at45dbxx_tx_rx_byte(0);
status = B_HAL_SPI_Receive(&_45DBXX_SPI, data, len);
B_HAL_GPIO_WritePin(_45DBXX_CS_GPIO, _45DBXX_CS_PIN, GPIO_PIN_SET);

return (status == HAL_OK);
}

Could be a CS timing issue but I doubt it because the device is very fast; much faster than the above code.

Title: Re: GCC compiler optimisation
Post by: gf on August 14, 2021, 10:00:52 am
Quote from: peter-h
Quote
The fact that it's on the stack means the compiler sees its entire lifespan, from function entry to exit, and knows it's never used.
Hmm. An interesting theory.

The primary purpose of a program is not to load values from memory locations, or to store values to memory locations. At the end, only the "visible side effect" count. How the visible effects are calculated does not matter. If the compiler can calculate the same visible side effects w/o accessing any memory, then it does not need to generate any load/store instructions. And if it can prove that a variable is completely unused (after eliminating the load/store instructions), then it can eliminate does not need to allocate any storage for the variable either.

Only volatile load/store operations are considered "visible side effect", i.e. to have a self-purpose, therefore they cannot be optimized out.

Mentally, it may be is better if you do not consider a variable a memory location in the first place, but rather consider it a name for a value (wherever the value happens to be stored, if stored at all).
It depends on the circumstances whether the compiler eventually allocates a memory location for a variable, or not.
It may also help to think in terms of SSA (https://en.wikipedia.org/wiki/Static_single_assignment_form) (in particular regarding local variables inside functions) which is the basis for virtually all optimzers today.

Quote
I like the "memcpy() doesn't know about volatile" explanation better, technically.

Of course. Even if memcpy() is a compiler intrisic, the compiler still has to handle an explicit call to memcpy() as if an inline funtion with the signature

Code: [Select]
void* memcpy(void * restrict dest, const void * restrict src, size_t count);

were called, and inlined into the code. I.e. any volatile qualifiers are lost when the arguments are passed to the function parameters dest and src.
And inside memcpy(), the load/store from/to source and destination buffer are NOT volatile operations, so they are allowed to be eliminated if they are needless for calculating the actual visible side effects.
Title: Re: GCC compiler optimisation
Post by: Siwastaja on August 14, 2021, 11:51:27 am
Mentally, it may be is better if you do not consider a variable a memory location in the first place, but rather consider it a name for a value (wherever the value happens to be stored, if stored at all).

This, this and this. You need to get rid of the "portable assembler" mindset and accept that C is a high level language. You describe the expected outcome of the program by sequential statements, then compiler is free to do anything to achieve the end-result your statements define. Because machine language is typically also sequential operations, you may incorrectly think these two sequences map directly or nearly directly to each other, but they don't (except by accident). It's also incorrect way of thinking that there "should" be direct mapping and it's just the "optimizer" that messes this up.

C does have some low-level features (like the volatile qualifier) that make it usable for near-hardware system programming, but the whole language doesn't work that way like people tend to expect.
Title: Re: GCC compiler optimisation
Post by: DavidAlfa on August 14, 2021, 01:26:59 pm
B_HAL_GPIO_WritePin(_45DBXX_CS_GPIO, _45DBXX_CS_PIN, GPIO_PIN_RESET);

Probably you'll need to put a small delay after that. I always start with slow SPI clocks and adding big delays when toggling pins.
Then, when it's already working, I start optimizing things.
Otherwise you add a lot of factors hat could be causing the problem. Too fast clock? Too fast CS? Who knows!

If you don't have a logic analyzer yet, get one! You have cheap 8-ch 24MHz analyzers in ebay.
They're amazingly handy when troubleshooting digital communications.
You'll quickly find out if there's something wrong in the bus.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 14, 2021, 04:07:09 pm
I wish it was that... Tcss etc are a mere 5ns

(https://peter-ftp.co.uk/screenshots/202108144912830317.jpg)

I have a 500MHz DSO and this was checked.

The serial FLASH runs off SPI2 which runs at 21mbps (the max possible on 32F4 SPI2/3). The FLASH can do 85MHz.

A timing issue does nevertheless remain, and it would be a bastard to find.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 14, 2021, 05:14:55 pm
Quote
The fact that it's on the stack means the compiler sees its entire lifespan, from function entry to exit, and knows it's never used.
Hmm.  An interesting theory.
If I change the buffer to "static volatile", then gcc will produce code to do the copy.  But clang doesn't. :-)
I like the "memcpy() doesn't know about volatile" explanation better, technically.

As some of us said, you can't expect any particular behavior for those cases because they are entirely implementation-dependent. And yes, you can see some pretty weird stuff when playing with this.

When *writing* to objects on the stack that are never *read* afterwards before they get out of scope, the compiler is free to optimize out those writes entirely. From a purely functional POV, those writes would have absolutely ZERO effect. The compiler assumes that the local stack is under its full control - so in some cases, even with a volatile qualifier, the code may be optimized out if the local variables in question are never read.

Now if you're willing to "play" with the stack on a very low-level (for instance for writing stack sentinels or something), your best bet is probably to do that directly in assembly.
Title: Re: GCC compiler optimisation
Post by: Doc Daneeka on August 15, 2021, 05:03:48 am
Quote
When *writing* to objects on the stack that are never *read* afterwards before they get out of scope, the compiler is free to optimize out those writes entirely. From a purely functional POV, those writes would have absolutely ZERO effect. The compiler assumes that the local stack is under its full control - so in some cases, even with a volatile qualifier, the code may be optimized out if the local variables in question are never read

That's not right: the purpose of volatile is to define what is observable behaviour of the program: the concept of a stack is implementation, it doesn't matter what the scope of the variable is, if it's volatile all accesses to it have to made strictly per the semantics of C. As far as C language is concerned, there is no concept of stack just objects that can be accessed (read or write). Although its implementation defined what it means for an access to be 'obervable', it cannot 'optimise away' a volatile access, be declaring it volatile you are telling the compiler it is observable.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 15, 2021, 06:26:15 am
Does "static" have any effect on whether unused storage is optimised away?

What concerns me is what looks like cases where say you have a 512 byte array and you use bytes 0,1,2 for flags and never read (or explicitly write, a byte at a time) the rest; the compiler might optimise away operations which write the whole array.

That could screw up say a serial FLASH which is expecting to always get 512 bytes transferred.
Title: Re: GCC compiler optimisation
Post by: Siwastaja on August 15, 2021, 06:47:31 am
If you have a memory-mapped hardware device expecting 512 bytes to be transferred, then this is the most obvious textbook example where volatile qualifier is required.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 15, 2021, 08:10:34 am
OK; let's look at whether "pData" or  "len" could get optimised out here

Code: [Select]

HAL_StatusTypeDef B_HAL_SPI_Transmit(SPI_HandleTypeDef *hspi, uint8_t *pData, uint16_t Size)
{
//  uint32_t tickstart;
  HAL_StatusTypeDef errorcode = HAL_OK;
  uint16_t initial_TxXferCount;

  initial_TxXferCount = Size;

  if (hspi->State != HAL_SPI_STATE_READY)
  {
    errorcode = HAL_BUSY;
    goto error;
  }

  if ((pData == NULL) || (Size == 0U))
  {
    errorcode = HAL_ERROR;
    goto error;
  }

  /* Set the transaction information */
  hspi->State       = HAL_SPI_STATE_BUSY_TX;
  hspi->ErrorCode   = HAL_SPI_ERROR_NONE;
  hspi->pTxBuffPtr  = (uint8_t *)pData;
  hspi->TxXferSize  = Size;
  hspi->TxXferCount = Size;

  /*Init field not used in handle to zero */
  hspi->pRxBuffPtr  = (uint8_t *)NULL;
  hspi->RxXferSize  = 0U;
  hspi->RxXferCount = 0U;
  hspi->TxISR       = NULL;
  hspi->RxISR       = NULL;

  /* Configure communication direction : 1Line */
  if (hspi->Init.Direction == SPI_DIRECTION_1LINE)
  {
    SPI_1LINE_TX(hspi);
  }

  /* Check if the SPI is already enabled */
  if ((hspi->Instance->CR1 & SPI_CR1_SPE) != SPI_CR1_SPE)
  {
    /* Enable SPI peripheral */
    __HAL_SPI_ENABLE(hspi);
  }

  /* Transmit data in 8 Bit mode */

    if ((hspi->Init.Mode == SPI_MODE_SLAVE) || (initial_TxXferCount == 0x01U))
    {
      *((__IO uint8_t *)&hspi->Instance->DR) = (*hspi->pTxBuffPtr);
      hspi->pTxBuffPtr += sizeof(uint8_t);
      hspi->TxXferCount--;
    }
    while (hspi->TxXferCount > 0U)
    {
      /* Wait until TXE flag is set to send data */
      if (__HAL_SPI_GET_FLAG(hspi, SPI_FLAG_TXE))
      {
        *((__IO uint8_t *)&hspi->Instance->DR) = (*hspi->pTxBuffPtr);
        hspi->pTxBuffPtr += sizeof(uint8_t);
        hspi->TxXferCount--;
      }
    }


  /* Clear overrun flag in 2 Lines communication mode because received is not read */
  if (hspi->Init.Direction == SPI_DIRECTION_2LINES)
  {
    __HAL_SPI_CLEAR_OVRFLAG(hspi);
  }

  if (hspi->ErrorCode != HAL_SPI_ERROR_NONE)
  {
    errorcode = HAL_ERROR;
  }

error:
  hspi->State = HAL_SPI_STATE_READY;
  return errorcode;
}


I just can't see it. It would need to be tracking the code calling this function, and realising that of the 512 bytes which are always transferred, most (or even all) have not changed.

Should I use this function prototype

Code: [Select]
HAL_StatusTypeDef B_HAL_SPI_Transmit(SPI_HandleTypeDef *hspi, volatile uint8_t *pData, volatile uint16_t Size);
The most context-narrow failure I am getting is that Windows cannot format the drive, if -Og is used.

EDIT: this code uses 2k buffers for the USB transfers. I tried to make them volatile but it didn't change anything.

EDIT2: compiling just the low level 45dbxx stuff with -O0 makes formatting work fine. So one "optimisation vulnerability" was definitely in there. Can't find it though. Shall I offer 50 quid (Paypal) to anybody who can find it? :)
Title: Re: GCC compiler optimisation
Post by: gf on August 15, 2021, 09:57:17 am
Should I use this function prototype
Code: [Select]
HAL_StatusTypeDef B_HAL_SPI_Transmit(SPI_HandleTypeDef *hspi, volatile uint8_t *pData, volatile uint16_t Size);

pData is not dereferenced in the given code snipplet. Therefore it makes no difference whether it is declared uint8_t* or volatile uint8_t*.

The interesting statement  where volatile matters is rather this one:
Code: [Select]
      *((__IO uint8_t *)&hspi->Instance->DR) = (*hspi->pTxBuffPtr);

If __IO is a macro which expands to volatile, then the assignment to *((__IO uint8_t *)&hspi->Instance->DR) becomes a "visible side effect" and cannot be "optimized out". Consequently the value which is assigned also needs to be calculated (i.e. fetched from *hspi->pTxBuffPtr).
Title: Re: GCC compiler optimisation
Post by: peter-h on August 15, 2021, 10:23:18 am
Yes __IO is volatile.

DR is the SPI "UART" data register.

Code: [Select]
  /* Set the transaction information */
  hspi->State       = HAL_SPI_STATE_BUSY_TX;
  hspi->ErrorCode   = HAL_SPI_ERROR_NONE;
  hspi->pTxBuffPtr  = (uint8_t *)pData;
  hspi->TxXferSize  = Size;
  hspi->TxXferCount = Size;
Title: Re: GCC compiler optimisation
Post by: cv007 on August 15, 2021, 02:42:29 pm
>Could be a CS timing issue but I doubt it because the device is very fast; much faster than the above code.

Your device will have a datasheet and give the timing requirements for CS, so probably not a bad idea to figure out if timing requirements are being met. Probably not important when you are running a 24MHz mcu (you see low the ns timing requirement, and know you cannot possibly fail), but you are into higher cpu speeds where you may no longer be able to 'eyeball' it. There is also OSPEEDR for a gpio pin which may also come into play (default is LOW SPEED). I would assume you have a way to measure this, so measure it and see what you are getting- if its good in any optimization then check it off your list and look elsewhere.

I always use -Os from start to finish, so see the compiler generated asm in a consistent way and also get to deal with problems as I create them due to optimization. This also eliminates subtle timing problems that can show up even when your code is 'correct' and compiles at any optimization level (it should). Ideally, timing is also handled correctly no matter which optimization but is easy to forget when you are doing something that a peripheral is not taking care of (like CS)..

You can change optimization within a file (gcc), but probably not something you want to use except in special circumstances-

#pragma GCC push_options
#pragma GCC optimize ("-Os")
//code here
#pragma GCC push_options


Also, when working in -Os it sometimes can get a little difficult to pick out the generated asm code of interest so surrounding the code with nop's is one way to highlight it as mentioned before. I will sometimes make a function 'noinline' temporarily when I want a clearer view of it, then when I'm satisfied with what I see it will revert back to what is was and can be confident I'm getting the same thing although is now not as clear. Using an online compiler like godbolt.org is also a quick/good way to create/test code to see how the compiler acts, even though your mcu headers are not available.

Title: Re: GCC compiler optimisation
Post by: Doctorandus_P on August 15, 2021, 02:47:35 pm
7 pages in 2 weeks. Quite impressive, but did not read it all.
Code: [Select]
This is news to me; I thought that compilers didn't change the order of functions in a .c file :) Why should they?
A lot of processors have a limited range for relative adressing, or can use smaller pointers for smaller jumps, and this is one good reason to re-order the functions, and can avoid "trampolines"
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 15, 2021, 04:40:07 pm
Quote
When *writing* to objects on the stack that are never *read* afterwards before they get out of scope, the compiler is free to optimize out those writes entirely. From a purely functional POV, those writes would have absolutely ZERO effect. The compiler assumes that the local stack is under its full control - so in some cases, even with a volatile qualifier, the code may be optimized out if the local variables in question are never read

That's not right: the purpose of volatile is to define what is observable behaviour of the program: the concept of a stack is implementation, it doesn't matter what the scope of the variable is, if it's volatile all accesses to it have to made strictly per the semantics of C. As far as C language is concerned, there is no concept of stack just objects that can be accessed (read or write). Although its implementation defined what it means for an access to be 'obervable', it cannot 'optimise away' a volatile access, be declaring it volatile you are telling the compiler it is observable.

Sorry, but... no.... on almost all points you make here, except for the stack. The standard doesn't even talk about stacks, and any implementation is indeed free to implement local variables (and "return addresses") any way it sees fit. It just so happens that the most common way, by far, is to use stacks (the only C compiler that I vaguely remember of that didn't even use a stack was an old one for a tiny programmable chip, and it was not even really C-compliant anyway), which is why I used this as an example. But yeah, let's remove the "stack" term here, no problem with that. It was too specific.

The idea still remains: any *local* object that is not qualified static ceases to exist once it gets out of scope, so anything happening on such an object AFTER it goes out of scope just doesn't exist per the definition, wherever this object is actually stored (stack, registers, or whatever else the implementation does.)

As to volatile, it's - unfortunately - more subtle than what you said.
1. To begin with: you talk about "observable". The only sentence actually using this term (in C99 at least) in the std is in this part:

Quote
An object that is accessed through a restrict-qualified pointer has a special association
with that pointer. This association, defined in 6.7.3.1 below, requires that all accesses to
that object use, directly or indirectly, the value of that particular pointer.
The intended
use  of  the restrict qualifier  (like the register storage  class)  is  to  promote
optimization, and deleting all instances of the qualifier from all preprocessing translation
units  composing  a  conforming  program  does  not  change  its  meaning  (i.e.,  observable
behavior
).

From what I understand here, this defines an "observable behavior" as the *meaning* of a program. Problem here is: what is the meaning of a program? The way I get this is the same as what I meant by the "functional POV", so anything volatile-related, when it may have unknown side-effects, but no analyzable effect, is NOT observable behavior. I may be wrong here and I admit we are really nitpicking on terms. I could not find the definition of the "meaning of a program" in the std.

For instance, taking the typical "delay loop" example, is a "delay loop", doing absolutely nothing apart from taking CPU cycles, part of the meaning of the program? If you can answer this one by a resounding "yes", without a blink, and backing it up with solid arguments, you are better than I am.

2. More importantly, about the volatile qualifier: it's unfortunately more subtle than it looks. Let's again quote C99 for the relevant parts:

Quote
An  object  that  has  volatile-qualified  type  may  be  modified  in  ways  unknown  to  the
implementation or have other unknown side effects.  Therefore any expression referring
to such an object shall be evaluated strictly according to the rules of the abstract machine,
as described in 5.1.2.3.

So far so good. Looks like "volatile" will guarantee that such a qualified object is evaluated in all cases, right?
But we need to refer to the "rules of the abstract machine" it mentions. So, again, relevant parts:

Quote
Accessing a volatile object, modifying an object, modifying a file, or calling a function
that does any of those operations are all side effects,
which are changes in the state of
the  execution  environment.  Evaluation  of  an  expression  may  produce  side  effects.  At
certain specified points in the execution sequence called sequence points, all side effects
of previous evaluations shall be complete and no side effects of subsequent evaluations
shall have taken place. (A summary of the sequence points is given in annex C.)

Still looks, at this point, like the volatile object will be evaluated no matter what. But the following paragraph kind of ruins it all:

Quote
In the abstract machine, all expressions are evaluated as specified by the semantics. An
actual implementation  need  not  evaluate part of an expression if it can deduce that its
value is not used and that no needed side effects are produced (including any caused by
calling a function or accessing a volatile object)
.

So, it looks a bit like what I said earlier. Doesn't it? (See the part in bold.)
Title: Re: GCC compiler optimisation
Post by: ataradov on August 15, 2021, 05:00:06 pm
OK; let's look at whether "pData" or  "len" could get optimised out here
I think you are also confusing optimized out and code generated in a way that values are no longer traceable by the debugger.

If this code is optimized, then you would not see any transfers. If you can put a logic analyzer on the bus and see that there is a transfer and it is 512 bytes, then nothing was optimized.

Relying on the debugger for everything you do is a bad idea.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 15, 2021, 05:10:47 pm
Relying on the debugger for everything you do is a bad idea.

Definitely.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 15, 2021, 07:16:38 pm
Fixed the 45dbxx issue with -Og.

There was a command to bring the device out of deep powerdown, but it was never put into deep powerdown in the first place (the command to do that was never used) although the device needed 35us if you did use that command. It worked with -O0 but marginally.

I also found a number of other timing issues where -Og shortened some delays considerably. This was quite interesting. I did some of that stuff months ago, not realising the optimisation issues. A good way to achieve supposedly reliable "minimum" timing is by reading an IO port. It is defined as volatile so the IO read itself can't get optimised out, but the compiler probably removed everything else, and possibly inlined the IO reads :) The chip made it obvious: it is a display controller and the display got corrupted. It is also an exceptionally slow chip (max SPI clock is just 1MHz, with pretty long CS setup and hold times).

Whether this is "broken code" is debatable because these are standard practices in embedded devt, for decades. They just don't work on these chips and with a modern compiler.

So in the end it wasn't anything to do with volatile declarations because I still can't get my head around some of the issues where you might try to fill a 512 byte buffer but it doesn't actually happen because the compiler has worked out that the last 90% of it was never accessed afterwards :)

Ataradov's microsecond delay function has been very useful :)
Title: Re: GCC compiler optimisation
Post by: ataradov on August 15, 2021, 07:31:42 pm
Whether this is "broken code" is debatable because these are standard practices in embedded devt, for decades. They just don't work on these chips and with a modern compiler.
Simply because you have been doing something for decades and tools  were shit to really optimize anything, does not make it right.

Even your reliance on blocking loops for delays will break in the future once the hardware gets better.

I still can't get my head around some of the issues where you might try to fill a 512 byte buffer but it doesn't actually happen because the compiler has worked out that the last 90% of it was never accessed afterwards :)
This is a very rare case in full and complete  applications. It happens sometimes when you comment out some code for debugging, and compiler finds a way to eliminate more than you expected.
Title: Re: GCC compiler optimisation
Post by: ejeffrey on August 15, 2021, 07:48:31 pm

Quote
In the abstract machine, all expressions are evaluated as specified by the semantics. An
actual implementation  need  not  evaluate part of an expression if it can deduce that its
value is not used and that no needed side effects are produced (including any caused by
calling a function or accessing a volatile object)
.

So, it looks a bit like what I said earlier. Doesn't it? (See the part in bold.)

I think you are misreading that.  Calling a function or accessing a volatile object are examples of side effects that would prevent the implementation from omitting an expression.
Title: Re: GCC compiler optimisation
Post by: lucazader on August 15, 2021, 08:12:00 pm
Whether this is "broken code" is debatable because these are standard practices in embedded devt, for decades. They just don't work on these chips and with a modern compiler.

This is just completely not true. These "standard" embedded practices I see all the time result in hard to read code that is equally hard to maintain. It always seems to be held together by a shoestring and falls apart in a light breeze.
I really don't get why the embedded world has been so slow to learn from the rest of software development and adopt modern tools, compilers and best practices.

We use the latest gcc compiler at work with our STM32 based devices. Never have any issues with code breaking or getting optimised out, unless it was a mistake made by us in the code.
Especially don't have any issues with timing. But I guess thats because we don't implement our timers as blocking loops based on the a certain number of instructions.
The systick and the HAL_Delay that it is derived from is a super consistent way to get good reliable timing.
If I need anything that has more resolution than systick, for me the most reliable way is to start up a hardware timer (eg TIM6) as a 1uS tick rate.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 15, 2021, 08:52:17 pm

Quote
In the abstract machine, all expressions are evaluated as specified by the semantics. An
actual implementation  need  not  evaluate part of an expression if it can deduce that its
value is not used and that no needed side effects are produced (including any caused by
calling a function or accessing a volatile object)
.

So, it looks a bit like what I said earlier. Doesn't it? (See the part in bold.)

I think you are misreading that.

After re-reading it, possibly. The phrasing, admittedly, is not the best here. It's borderline confusing.

Now, try this piece of code. It's interesting:

Code: [Select]
#include <memory.h>

void my_memcpy(volatile void *dst, volatile void *src, size_t n)
{
    while (n--)
        *(char *)dst++ = *(char *)src++;
}

void f()
{
    volatile char buffer[512];
    my_memcpy(buffer, (volatile char *)0x08000000, 512);
}

In your opinion, should it or should it not do anything?
Title: Re: GCC compiler optimisation
Post by: gf on August 15, 2021, 09:40:15 pm
A bit questionable is certainly the cast from volatile void* to char*, as it drops volatile. I would rather not consider the memory access to *(char*)dest a volatile access any more.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 15, 2021, 10:11:02 pm
Should debugging be at all possible with anything other than -O0 or -Og?

I have tried it with -O3 and while stepping does work, sort of, it does bizzare things.

I wonder what the explanation for this is. Stepping works by temporarily inserting a special opcode in that place, although the 32F has a small number of dedicated hardware breakpoints which, when hit, substitute the instruction in place of the opcode fetch so that works on code in FLASH also.

"It always seems to be held together by a shoestring and falls apart in a light breeze."

You are missing the point. If you want to achieve a delay of say 100ns (min) from CS=0 to something happening, you aren't going to make a call to RTOS :) This is done by a short delay made up of non-removable instructions, and I don't think there is any other way to do it, short of external hardware which inserts wait states (which would be ridiculous in this case).

Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 15, 2021, 10:11:25 pm
A bit questionable is certainly the cast from volatile void* to char*, as it drops volatile. I would rather not consider the memory access to *(char*)dest a volatile access any more.

Alright, that's well spotted. Changing this to:
Code: [Select]
*(volatile char *)dst++ = *(volatile char *)src++;
does generate code for the copy.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 15, 2021, 10:17:39 pm
Should debugging be at all possible with anything other than -O0 or -Og?

I have tried it with -O3 and while stepping does work, sort of, it does bizzare things.

With optimizations, code may get simplified in ways you don't expect. Apart from code possibly being optimized out entirely, and local variables put in registers, one thing that frequently happens too is that code can get reordered compared to what you wrote in C, so that matching between source code and assembly code statement by statement is not guaranteed. So the debugger may appear to be jumping all over the place when single stepping. You often have to switch to assembly code view in this case to figure out what happens.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 15, 2021, 10:23:45 pm
Should debugging be at all possible with anything other than -O0 or -Og?
Yes, you just need to get used to not seeing everything and just filling the gaps in your head. Why do you need to see absolutely everything? If you really need that, then stepping though assembly is a better option anyway. 


Stepping works by temporarily inserting a special opcode in that place, although the 32F has a small number of dedicated hardware breakpoints which, when hit, substitute the instruction in place of the opcode fetch so that works on code in FLASH also.

Hardware breakpoints do not need any substitutions. They just stop execution when address comparator tells so. And placing transparently soft breakpoints in the flash is the worst debugger feature ever invented. Thankfully most sane debuggers let you disable that behaviour.
Title: Re: GCC compiler optimisation
Post by: lucazader on August 15, 2021, 11:06:46 pm
Should debugging be at all possible with anything other than -O0 or -Og?

I almost exclusively use -Os most of the time, including in debug builds.
As others have said you just have to get used to single stepping jumping around a bit.
I find it helps to also have logs as well as the debugger to decode what is going on.
Title: Re: GCC compiler optimisation
Post by: ejeffrey on August 16, 2021, 05:47:47 am
Should debugging be at all possible with anything other than -O0 or -Og?

I have tried it with -O3 and while stepping does work, sort of, it does bizzare things.

I wonder what the explanation for this is. Stepping works by temporarily inserting a special opcode in that place, although the 32F has a small number of dedicated hardware breakpoints which, when hit, substitute the instruction in place of the opcode fetch so that works on code in FLASH also.

Single stepping doesn't use explicit breakpoints at all, hardware or software.  It is a separate operation mode of the CPU that causes it to halt before every instruction.  The debugger then reads the PC and uses the debug symbols to map that address back to the source line of code that generated it.

The reason it jumps around is simply that the compiler re-ordered the code.

Title: Re: GCC compiler optimisation
Post by: DavidAlfa on August 16, 2021, 07:57:08 am
You should debug with O0 unless you want weird things happenning.
The compiler will optimize the code, re-use parts from other functions, update the variables on a different order or directly omitting them... you will find yourself in a no sense behaviour when debugging.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 16, 2021, 08:41:18 am
"You should debug with O0 unless you want weird things happenning.
The compiler will optimize the code, re-use parts from other functions, update the variables on a different order or directly omitting them... you will find yourself in a no sense behaviour when debugging."

I have come to the same conclusion.

Develop mostly with -Og (or whatever you fancy) and recompile with -O0 if you want to step through the code to check the logic.

But weird stuff seems to be going on anyway. Last night the whole thing was running with -Og. This morning I did some edits (to an unrelated file) and it no longer works unless the serial FLASH code is compiled with -O0 (which was the only way to make it work previously). Tried project/clean etc; the usual Cube IDE stuff one does when anything isn't working right. The serial FLASH code is 3rd party code and while it all seems right, something isn't right and I can't find it, and when I step through it (having compiled it with -Og) a lot of weird stuff is going on. One could tear out one's hair with this. That said, with that one file done with -O0, the project is rock solid.

Fairly obviously the only way to debug optimised code is with "printf" statements. I have a printf() which comes out via the SWV ITM console, and I also did a much more compact itm_puts() which does the same, and with ITM running at 2MHz it doesn't slow down the code much.

I reckon lots of people just work with -O0 and don't bother to ever change it, because everything works as you would expect it (or not) and you can immediately debug
the code if needed. There is a sizeable code penalty though: 220k versus 160k. That will not matter in most projects...

The problem is that not everybody is an expert, and e.g. -O3 opens up various traps. I was seeing strange stuff e.g. conditionals totally skipped. The test was checking a byte in a buffer which was read from an eeprom, but I could not easily examine the buffer since it was optimised out. By the looks of it, the compiler didn't implement the eeprom read loop.
Title: Re: GCC compiler optimisation
Post by: DavidAlfa on August 16, 2021, 08:55:28 am
Og does some optimizations, so it'll still mess things up, like invisible functions because they are inlined.

If things stop working when enabling optimizations, your code is wrong.
Most often it's because you're forgetting to declare volatiles when you should.
That also includes simple delay loops that increase a value, the compiler will see that it does nothing but losing time and remove them, so you'll need to declare the variable as as volatile or insert an asm("nop") in the loop which is also considered volatile and thus not optimized.
Title: Re: GCC compiler optimisation
Post by: westfw on August 16, 2021, 09:00:46 am
Quote
Calling a function or accessing a volatile object are examples of side effects that would prevent the implementation from omitting an expression.
So I would have believed.But apparently nowadays, if you insert a function call, it's getting harder and harder to tell whether that will actually result in an optimization-defeating function call.  It could get inlined by the compiler.  It could get inlined by the linker.  The compiler can decide that it knows what the function does and do it some other way that doesn't defeat optimization.

Is there a way to defeat this on a per-function-call basis?
Code: [Select]
  NOINLINE memcpy(p1, p2, l);
There claims to be an attribute:
Code: [Select]
void foo() __attribute__((no_builtin("memcpy"))) {   :   memcpy(b1, b2, l);}But it didn't seem to work in our example case.  (perhaps it still knew what memcpy() was supposed to do, and decided that was pointless, even if it would have avoided generating inline code if it hadn't been pointless?)
Title: Re: GCC compiler optimisation
Post by: peter-h on August 16, 2021, 11:09:22 am
"If things stop working when enabling optimizations, your code is wrong.
Most often it's because you're forgetting to declare volatiles when you should."

My 50 quid offer stands :) EDIT: NO LONGER I found the problem. It was another timing issue; CS was being raised while the last byte being written was still shifting into the device. It worked previously due to bloated ST SPI code, which in any case needed the 1kHz tick running, and here I had removed all that timeout stuff so it all ran a lot faster, but still not fast enough in -O0 to break it.

"That also includes simple delay loops that increase a value, the compiler will see that it does nothing but losing time and remove them, so you'll need to declare the variable as as volatile or insert an asm("nop") in the loop which is also considered volatile and thus not optimized."

All that stuff is gone. There are ~5us and similar delays which are implemented exactly as suggested, and checked with a DSO.

I've written masses of code which all runs fine. This issue seems to be in the serial FLASH code.

Perhaps I am not fully understanding what "volatile" does.

If you read say a UART register (which is #defined with __IO which is "volatile") and load that data into a 512 byte buffer, that buffer cannot be optimised out - because the data source is volatile. The whole buffer should not be optimised out even if you loaded only 3 bytes into it. Is that correct?

BTW I am told that for the last 20 years memcpy and memmove are the same, and memcpy performs a check for overlapping blocks.
Title: Re: GCC compiler optimisation
Post by: DavidAlfa on August 16, 2021, 12:12:49 pm
Accessing a peripheral is volatile.
Are your using you own function based on the HAL libraries, rigth? Try using HAL  to ensure there isn't a mistake somewhere.
I don't think you need volatile for anything unless written by an interrupt or because you specifically want to be able to check that variable when debugging optimized code. (Or the said "dumb" loops).

How does it exactly fail? Is something working at all? Try doing the most simple thing, reading the jedec/device id from it.
If even that fails, you have to find out the problem source. Of course, not easy because it only fails when optimizing.
You can specifically force the compiler to not optimize a function by adding this:
Code: [Select]
__attribute__((optimize("O0"))) void notOptimizedFunction(void){
}

My approach would be this:
- Disable optimizations for all flash-related functions. If it still doesn't work, your problem is elsewhere.
- If it works, optimize the functions, one at a time, until you find the one causing trouble.
Title: Re: GCC compiler optimisation
Post by: Nominal Animal on August 16, 2021, 01:47:07 pm
For the last twenty years, I've written a lot of C code, the vast majority compiled with GCC, some with Clang, Intel CC, Pathscale, or Portland Group compilers.  Always with -O2.  My bug rate is lower than average, too; I do claim to know this stuff.

(I'm also proficient in several other programming languages from Fortran to Python, and have a pretty widely ranging background in development, so I'm not "stuck" in C, or in imperative programming languages, either.  By this I mean, I have years of experience in software development in very different fields from web development to microcontrollers; I'm not a one-niche guy, who assumes their experience in one niche is extensible to everywhere else.  I've just found that in all niches where I've used C, -O2 has been the proper optimization choice.)

On x86 and x86-64, I do a lot of parallel and distributed processing.  For efficiency, I often use lockless structures via compiler-provided atomic built-ins.  A decade ago, I used to write a lot if extended inline assembly for SIMD operations, but nowadays the x86 intrinsics are so well integrated to the compilers it is no longer necessary.  So, I do claim I know quite a bit about complex interactions between threads, atomicity, and manipulating volatile data (pun intended).

The first MCU I started developing on was a Teensy 2.0++, an Atmel AT90USB1286, on top of a set of header files, avr-libc, and avr-gcc.
I have a few dev boards using ATtinys (digispark clones), ATmega32u4 (pro micro clones), and ATmega328 (pro mini clones) that I've programmed the same way; on bare metal.  For more interesting stuff, I now use ARM microcontrollers, in Arduino or PlatformIO environments.  (I particularly like Teensy LC, 3.2, 4.0, and 4.1, all of which I have at least one.  I do have about a dozen others, from various manufacturers, some still in their original packaging.)

On the electronics side, I'm an utter ham-handed hobbyist.  I do have a physics background, with theoretical courses over electronics (up to opamps and digital logic), as my "core field" is computational materials physics, specifically simulator software development, but have only used this "in anger" in the last few years, mostly using EasyEda and JLCPCB (because it's so darned easy there).  So, I'm still learning myself, and not "stuck" believing I know everything I need to know; I know I don't know enough, and am very interested in learning, and not at all afraid of admitting publicly when I'm wrong.  I'm deliberately very blunt that way.  It's cathartic, too.

GCC atomic built-ins (https://gcc.gnu.org/onlinedocs/gcc/_005f_005fatomic-Builtins.html) are available for ARM architectures, and are also provided by LLVM Clang (ie. if you use clang to compile to ARM targets).  Do not let the C++18 reference mislead you; all they mean is that the six __ATOMIC_ memory order constraints use the memory model definitions in the C++18 standard, that's all.  It is perfectly acceptable to use these in plain C code, or embedded C++.  I typically end up using __ATOMIC_SEQ_CST anyway.  The "trick" is to always use the atomic built-in when accessing a variable, and not mix non-atomic accesses with atomic ones, unless the non-atomic accesses are allowed to occasionally be garbled (like in a compare-and-swap loop).

The C standard is defined in terms of an abstract machine.  Mostly, C compilers strictly follow this standard.  GCC (and other compilers) have options that diverge or relax some rules.  Generally, -O2 does not include any of those.  You can check by examining the output of gcc -c -Q -O2 --help=optimizers.
The C standard does leave many things "implementation defined".  (When you use freestanding C++, for example in the Arduino environment, almost everything is "implementation defined" per the C++ standard; only when you know a feature is available, can you reasonably consult the C++ standard to see how it is supposed to be implemented.  The corresponding freestanding C environment is much more "defined", so that the environment normally used for microcontroller development is quite a complicated subset of C and C++.)

The most relevant terms here are immutable, constant, const, volatile, and atomic.

Atomic is the most complicated one, because it really refers to several things that attempt to achieve the same result.  In certain architectures, basic accesses to base types may be inherently atomic; this depends on the hardware architecture. Later C and C++ standards define atomic types, corresponding to these types.  GCC and LLVM-clang provide the aforementioned builtins, that implements such atomic accesses, if it is possible in the current hardware; if the target is such an inherently atomic type, these builtins compile to basic accesses, and thus are very efficient ways to implement atomic accesses.  In both cases, the problem is that not all hardware architectures implement the full complement –– in particular, most are either compare-exchange or load-locked, store-conditional type –– so one kinda-sorta needs to check the compiler generated assembly or machine code to see if the constructs you need generate sane-looking code.  If there are things like "disable interrupts", you know that operation isn't really atomic on that architecture, and the compiler is trying to work around the hardware; a different code pattern is then needed on that hardware.

Constant != const.  The C and C++ standards use specific definitions for terms like "literal constant" and "constant"; they are not necessarily what one might think they mean.  So, for constant, be careful to check the context in which it is used.  (Also, if you find your compiler does not do something that the standard says it should, means either the compiler has a bug, or the compiler developers and you see that passage in the standard differently.  I always say that reality trumps theory, because it does.  Instead of railing against it, it is more effective to report it (but accept that it likely will be ignored) and work around it, because what matters is that the generated code works as required in all situations in real life; whether it is exactly according to rules drawn up by a committee is always a secondary concern, something for the business and people staff to discuss in their endless meetings.)

Immutable is used in the sense that "this is not allowed to be modified".  From the C programmers view, an immutable object or variable resides in read-only memory, and an attempt to modify such causes "undefined behaviour", something that depends on the hardware and the environment used.  In userspace code running under a full operating system, it usually leads to segmentation violation error, and a crash of that process.  In a microcontroller, the attempt may be ignored, the MCU can reset, or an interrupt fire.  It varies.

This leaves the two C keywords, const and volatile.  const is a promise from the programmer to the compiler that the code does not try to modify the object or variable such denoted.  volatile is the inverse: it means the compiler is not allowed to make any assumptions whatsoever about the object or variable such denoted.

This means that constructs such as const volatile int  foo; are perfectly valid and useful.  The const keyword is a promise to the compiler that the code in this scope will not try to modify foo, and the volatile keyword tells the compiler that whenever foo is used, it must read its value from memory, because it may be modified by something unknown; even by hardware, another thread, whatever.

Trick is, const and volatile work exactly that way in an expression as well.  Even if you have an object or variable not declared volatile, you can take its address, cast that to a pointer to volatile to the type of the object/variable, and dereference the cast; such access is then equivalent to one when the variable or object was declared volatile in the first place.  I do not recommend this as a general pattern, because it means the type of that variable must be duplicated in every such cast.
A much better pattern is to have the variable or object declared volatile, but in any scope where a snapshot of that suffices, just copy its value into a local const one (non-volatile).
 
There are a couple of additional details in C that are useful when dealing with numerical expressions (important when you get into stuff like Kahan summation), without going into the details of how that abstract machine works and what side effects and sequence points are.  First is that in C99 and later, casts of numeric types, both reals and integers, limit the range and precision to that of the cast type.  The second is that unsigned integer arithmetic is modulo arithmetic (wraps around), and since C99, there are exact-width unsigned binary types uintN_t, binary twos complement types intN_t (and corresponding minimum-width and optimum-width/fast types) provided by the compiler in <stdint.h> even in freestanding environments (i.e., always).  As of 2021-08-16, the fixed-point support in GCC is not good enough to really use in my opinion.  Using the integer overflow built-ins (https://gcc.gnu.org/onlinedocs/gcc/Integer-Overflow-Builtins.html#Integer-Overflow-Builtins) you can do multi-limb (multi-byte/word) counters trivially.  (Making one atomic really needs two generation counters, and a retry loop, though; and for both reading and incrementing/modification.)

Finally, compiler barriers, in particular __asm__ __volatile__("": : :"memory"); , can be used to ensure all memory accesses in preceding code are done prior to this barrier, and all memory accesses done in succeeding code are done after this barrier.  Basically, it makes sure the compiler does not move memory accesses across this barrier.  (It does this by basically telling the compiler that everything it knows about memory contents at this exact point becomes invalid.)

I really, really do not understand why you'd find -O0 necessary.  I suspect it is because you haven't really yet grokked how C compilers and the C language work, deep inside the nitty gritty details.  This is not an insult; not all C programmers need that kind of deep understanding to effectively wield C in anger, but since you found you need to disable optimizations to get the code you want, I suspect you do need to know.  I warmly recommend reading the standard; specifically, starting with the C99 version, because it is the most widely supported one (except for Microsoft C++ compiler, as Microsoft still refuses to fully support C99, even after contributing significantly to later C11 version).  The final draft, with the three corrigenda included, is publicly available as n1256.pdf (http://www.open-std.org/jtc1/sc22/WG14/www/docs/n1256.pdf) at open-std.org.  (For C11, the final draft is n1570.pdf (http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1570.pdf), and for C18, archived as n2176.pdf (https://web.archive.org/web/20181230041359/http://www.open-std.org/jtc1/sc22/wg14/www/abq/c17_updated_proposed_fdis.pdf).)  The actual standards can be bought from ISO, but I haven't bothered; too expensive for what they are.
Title: Re: GCC compiler optimisation
Post by: Bassman59 on August 16, 2021, 03:21:40 pm
My 50 quid offer stands :) EDIT: NO LONGER I found the problem. It was another timing issue; CS was being raised while the last byte being written was still shifting into the device.

I hate when that happens!

Title: Re: GCC compiler optimisation
Post by: peter-h on August 16, 2021, 03:36:46 pm
Everything now works with -Og.

However, -O3 still breaks things. I will work on it using the __attribute__((optimize("O0"))) void notOptimizedFunction(void) - thank you for that tip.

It is just difficult to debug -O3 code. But it does look like it is all SPI related. For example neither myself nor my colleague can see the reason for the special treatment of the 1 byte transfer case (SPI_MODE_SLAVE is false, btw):

Code: [Select]
  /* Transmit and Receive data in 8 Bit mode */

    // The need for this initial byte is unknown
    if ((hspi->Init.Mode == SPI_MODE_SLAVE) || (initial_TxXferCount == 0x01U))
    {
      *((__IO uint8_t *)&hspi->Instance->DR) = (*hspi->pTxBuffPtr);
      hspi->pTxBuffPtr++;
      hspi->TxXferCount--;
    }

    while ((hspi->TxXferCount > 0U) || (hspi->RxXferCount > 0U))
    {
      /* Check TXE flag */
      if ((__HAL_SPI_GET_FLAG(hspi, SPI_FLAG_TXE)) && (hspi->TxXferCount > 0U) && (txallowed == 1U))
      {
        *(__IO uint8_t *)&hspi->Instance->DR = (*hspi->pTxBuffPtr);
        hspi->pTxBuffPtr++;
        hspi->TxXferCount--;
        /* Next Data is a reception (Rx). Tx not allowed */
        txallowed = 0U;
      }

      /* Wait until RXNE flag is reset */
      if ((__HAL_SPI_GET_FLAG(hspi, SPI_FLAG_RXNE)) && (hspi->RxXferCount > 0U))
      {
        (*(uint8_t *)hspi->pRxBuffPtr) = hspi->Instance->DR;
        hspi->pRxBuffPtr++;
        hspi->RxXferCount--;
        /* Next Data is a Transmission (Tx). Tx is allowed */
        txallowed = 1U;
      }
    }

but sure as hell if you don't do that, it doesn't work. And it isn't a case of needing to transmit a byte to start things off (like you have to do with interrupt-driven UART TX usage) because what if you transmitted 2 bytes?

In the function which just transmits data (with SPI you always receive the same # as you transmit, but in this case the RX is discarded) I have the test at the end before CS is raised and that was a major bug which didn't surface with slower code

Code: [Select]
    /* Transmit data in 8 Bit mode */

  // The need for this initial byte is unknown
    if ((hspi->Init.Mode == SPI_MODE_SLAVE) || (initial_TxXferCount == 0x01U))
    {
      *((__IO uint8_t *)&hspi->Instance->DR) = (*hspi->pTxBuffPtr);
      hspi->pTxBuffPtr++;
      hspi->TxXferCount--;
    }

    while (hspi->TxXferCount > 0U)
    {
      /* Wait until TXE flag is set before loading data */
      if (__HAL_SPI_GET_FLAG(hspi, SPI_FLAG_TXE))
      {
        *((__IO uint8_t *)&hspi->Instance->DR) = (*hspi->pTxBuffPtr);
        hspi->pTxBuffPtr++;
        hspi->TxXferCount--;
      }
    }

    // wait for last byte to get shifted out - needed before CS is raised!
    while ( !__HAL_SPI_GET_FLAG(hspi, SPI_FLAG_TXE) ) {}
Title: Re: GCC compiler optimisation
Post by: DavidAlfa on August 16, 2021, 03:51:44 pm
Oh, were you reading the buffer empty flag instead the shift register?
That's a typical mistake, as the spi peripheral is double buffered!
These are the 3 lines of code that take 80% of developing time :D
By the way, O3 barely makes any difference vs O2, but takes more space.
You also have Ofast if you prefer speed over code size, the performance boost is noticeable, at least with the stm32.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 16, 2021, 04:24:25 pm
Ha! It gets better :)

I made the same mistake which as you say so many have made. One has to wait for TX buffer empty AND TX shift register empty, before raising the slave device CS to 1.

Neither was happening and my above line is doing only half of it. This is needed:

Code: [Select]
     // wait for last byte to get shifted out - needed before CS is raised!
    while ( (__HAL_SPI_GET_FLAG(hspi, SPI_FLAG_TXE)==1) && (__HAL_SPI_GET_FLAG(hspi, SPI_FLAG_BSY)==0) ) {}

It was working by accident, at the combination of 21mbps and the bloated code.

The 32F4 does not have a "TX shift reg empty" bit as such. It has SPI_FLAG_BSY which looking at some original ST code does that job, but checking both bits is better in case you catch it just as the byte is being transferred from the holding reg to the shift reg.

This is an issue only on HAL_SPI_Transmit; on HAL_SPI_Transmit_Receive the code waits for the incoming byte to fully arrive so this is not an issue. Also I notice the original ST code contains, via a convoluted route involving system tick based timeouts, and of course a lot of code, the above test for SPI_FLAG_BSY... That is not great because in many applications you will not want a blocking function for SPI TX because it prevents you doing something useful while the last byte is going out, and doing osDelay(1) (to yield to the RTOS) is really crude because it will kill the data rate.

So far, not had issues apparently related to "volatile". It isn't trivial to try that because if you have a buffer and call some function to fill that, and you put "volatile" in front of that buffer, it complains that the volatile was discarded for the function. So you have to change a lot of stuff.

Now running with -O3 and no hacks.
Title: Re: GCC compiler optimisation
Post by: JOEBOBSICLE on August 16, 2021, 04:37:55 pm
Why don't you just use the HAL?
Title: Re: GCC compiler optimisation
Post by: peter-h on August 16, 2021, 04:41:14 pm
The HAL SPI functions, unmodified, need interrupts to implement timeouts. I am doing a boot loader which has these disabled.
Title: Re: GCC compiler optimisation
Post by: JOEBOBSICLE on August 16, 2021, 04:54:05 pm
Seems overly complicated instead of replacing whatever the st function they use to read the tick for timeouts.

Changing one function (HAL_GetTick())  VS all the places they implement a timeout. Surely it just makes sense to do a small change?
Title: Re: GCC compiler optimisation
Post by: DavidAlfa on August 16, 2021, 05:47:37 pm
Why can't you run the systick timer in the bootloader?
If you want to momentary disable interrupts, use __disable_irq() and __enable_irq() macros.

Systick doesn't cause any harm, it just increases the variable "uwTick" and returns.

It'll probably work anyways with the timer disabled.
The timeouts will never work, so it'll stall forever if something fails.
But as you said, that's impossible to happen in your application.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 16, 2021, 06:20:21 pm
Yes; very good points.

I did not want interrupts because setting up the vectors to movable ISRs is an extra complication. There were other reasons too; for example I don't like to throw in a ton of code which is so bloated that I can hardly read it. Remember I am new to the 32F world, coming from much simpler CPUs like the H8, and programming these mostly in assembler where everything does exactly what you want it to do (well, usually) :) I want to simply stuff to a point where I can understand it fully because in the long run I will be supporting the product alone. In my little business we still sell products I designed in the 1990s.

I am learning but at a finite rate, and it is a steep curve, with the complicated Cube IDE animal to deal with too :)

When designing the PCB I read the 300 page hardware manual word for word and the PCB worked first time with no bugs. But I didn't read the 2000 page Reference Manual in the same way. Only when needed, and then one realises that 300 lines of ST HAL code is really about 10 lines if you write code to do the job you actually need. And their code can hide a lot of problems, and it does - the bugginess of HAL code is well known. So another reason to simplify the HAL code.

On the topic of optimisation, this is interesting: With -O3, the compiler replaces this

Code: [Select]
for (uint32_t i=0; i<length; i++)
{
buf[offset+i]=data[i];
}

with memcpy() and guess what happens when this code runs in the boot loader? The thing bombs, because mmcpy is way high up in the FLASH and well outside the boot loader! This was discovered only because Cube shows a stack trace and on there appeared memcpy which does not exist anywhere near there.

And people here tell me that optimised code which doesn't run is "broken" :)

Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 16, 2021, 06:45:06 pm
Why bother with this "systick" interrupt when you're at a low level stage?
Just use the DWT cycle counter, as is often suggested in this forum. It just needs to be enabled. I use this for enabling it in C:
Code: [Select]
void DWT_Init(void)
{
if (! (CoreDebug->DEMCR & CoreDebug_DEMCR_TRCENA_Msk))
CoreDebug->DEMCR |= CoreDebug_DEMCR_TRCENA_Msk;

DWT->CYCCNT = 0;
DWT->CTRL |= DWT_CTRL_CYCCNTENA_Msk;
}

Enabling it in assembly is also just a couple instructions if you need to do this in the startup code or something.

Then you can read the DWT->CYCCNT register. Resolution is the system clock period. Doesn't require any interrupt or any peripheral.

Title: Re: GCC compiler optimisation
Post by: JOEBOBSICLE on August 16, 2021, 07:40:41 pm
Yes; very good points.

I did not want interrupts because setting up the vectors to movable ISRs is an extra complication. There were other reasons too; for example I don't like to throw in a ton of code which is so bloated that I can hardly read it. Remember I am new to the 32F world, coming from much simpler CPUs like the H8, and programming these mostly in assembler where everything does exactly what you want it to do (well, usually) :) I want to simply stuff to a point where I can understand it fully because in the long run I will be supporting the product alone. In my little business we still sell products I designed in the 1990s.

I am learning but at a finite rate, and it is a steep curve, with the complicated Cube IDE animal to deal with too :)

When designing the PCB I read the 300 page hardware manual word for word and the PCB worked first time with no bugs. But I didn't read the 2000 page Reference Manual in the same way. Only when needed, and then one realises that 300 lines of ST HAL code is really about 10 lines if you write code to do the job you actually need. And their code can hide a lot of problems, and it does - the bugginess of HAL code is well known. So another reason to simplify the HAL code.

On the topic of optimisation, this is interesting: With -O3, the compiler replaces this

Code: [Select]
for (uint32_t i=0; i<length; i++)
{
buf[offset+i]=data[i];
}

with memcpy() and guess what happens when this code runs in the boot loader? The thing bombs, because mmcpy is way high up in the FLASH and well outside the boot loader! This was discovered only because Cube shows a stack trace and on there appeared memcpy which does not exist anywhere near there.

And people here tell me that optimised code which doesn't run is "broken" :)

You're trying to run a bootloader out of ram which is somewhere in your flash image right? I have written 4 different bootloaders for STM chips and just allocate a section of flash for the bootloader and jump to the application code after X amounts of seconds since boot. Running out of ram is annoyingly complicated and dangerous imo. What if you flash a hex that's dangerous?

If you want code to last decades then consider how easy it'll be to port to a new chip. Your current chip may only last another decade. You'll likely need to rip up all your drivers so it's worth making them abstract and easy to port.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 16, 2021, 08:01:37 pm
The thing bombs, because mmcpy is way high up in the FLASH and well outside the boot loader!
This should not happen. How did it happen? A bootloader is a self contained project. It either builds standalone and works or does not.

And people here tell me that optimised code which doesn't run is "broken" :)
You are doing something highly unconventional and questionable, so of course it does not work.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 16, 2021, 08:05:09 pm
Running out of ram is annoyingly complicated and dangerous imo. What if you flash a hex that's dangerous?
All the bootloaders I've written for ARM run from RAM. At least the ones that would fit into RAM.

Sure, there is danger associated with updating the bootloader itself, but in other case there is not even an option. And having option and not using them is nicer than not having options.

There is nothing inherently dangerous with running from RAM. It also lets me receive the next packet of data while flash is busy writing the previous one, instead of execution just being blocked.
Title: Re: GCC compiler optimisation
Post by: Nominal Animal on August 16, 2021, 08:08:04 pm
If you do examine the size of the generated code, you'll find that -O3 significantly expands the size of the code generated.  Because of that, -O3 is rarely used.  You really should consider using -O2 instead; and only if necessary, add specific additional optimization flags.  (I occasionally use -fno-trapping-math -ffinite-math-only, for example.)

On the topic of optimisation, this is interesting: With -O3, the compiler replaces this [...code..] with memcpy()
Yes, because a compiler is allowed to turn code into a (standard) function call, when that standard function call is available, and -O3 enables all sorts of "unsafe" optimizations and assumptions the compiler may do to get stuff "run faster".  Usually it fails at it, though, so -O3 is rarely used.

If you had read the GCC Standards conformance (https://gcc.gnu.org/onlinedocs/gcc-11.2.0/gcc/Standards.html#C-Language) statement, second to last paragraph in section 2.1: "GCC requires the freestanding environment provide memcpy, memmove, memset and memcmp", you'd have known these four functions need special care because of this, when dealing with such a strictly constrained situation as you have right now.

GCC has built-ins (https://gcc.gnu.org/onlinedocs/gcc/Other-Builtins.html) for a lot of functions in the C library, but the above four are the ones that need to be carefully handled in a very strictly constrained situation like yours, because GCC may emit calls to them at any point.  Its documentation says so.

This is not rare, either; just normally unobserved.  For example, it is the GCC built-in printf() that turns printf("Hello, world!\n") into a puts("Hello, world!") or fputs("Hello, world!\n", stdout); library call.  It is only because you have a very constrained situation that you noticed.
(Then again, very few programmers seem to know about this.)

These are not difficult to fix, although the optimum way always depends on the situation.  I like to implement my own functions (with a prefix to the name, say local_, and do explicit calls to those).  Usually the header file describing them has GCC preprocessor magic (including __builtin_constant_p(), for other but similar functions _Generic() too to do different things depending on the type of the arguments) so that when the targets are known to be aligned or of specific types, aligned/optimized versions of the functions are used.

You can trivially check after compilation but before linking using e.g.
    LANG=C LC_ALL=C objdump -tT object-file 2>/dev/null | awk '$2 == "*UND*" { print $4 }'
to list all the symbols object-file uses, but does not define itself.  In your case, when run against the object file containing your RAM-based Flash updater functions, it would have listed memcpy.

In script form,
    #!/bin/sh
    export LANG=C LC_ALL=C
    status=0
    for obj in "$@" ; do
        if objdump -tT "$obj" 2>/dev/null | awk -v obj="$obj" 'BEGIN { rc=0 } $2 == "*UND*" { printf "%s: References external symbol: %s\n", obj, $4 ; rc=1 } END { exit rc }' ; then
            status=1
        fi
    done
    exit $status
this can be used in any Makefile.  If it lists any symbols, it will return failure.  In a Makefile, if your rule for building this object file contains, after the compile command ($(CC) $(CFLAGS) -c $^), a call to this script naming this file ($(NOEXTSYMS) $@ where NOEXTSYMS is a Makefile variable defining the path to the above script), the build will fail if such external symbols are found in the object file.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 16, 2021, 08:40:27 pm
"Just use the DWT cycle counter"

Thanks for the tip. I didn't know about that. However, the fact is that those error conditions are impossible unless the silicon is defective, so why carry so much bloat to pick up conditions which a) will never be seen and b) cannot be reported to the user. Example: I've sold, since 1995, about 10k of a box with an H8/300, 62256 SRAM, 28C256 EEPROM. The number of actual hardware failures could be counted in the fingers of one hand.

"Yes, because a compiler is allowed to turn code into a (standard) function call, when that standard function call is available, and -O3 enables all sorts of "unsafe" optimizations and assumptions the compiler may do to get stuff "run faster".  Usually it fails at it, though, so -O3 is rarely used."

That isn't exactly what those above have been telling me (that any failure due to optimisation means my code is crap) :) I can see this "optimisation" is legitimate, but that's not the same Q.

"All the bootloaders I've written for ARM run from RAM. At least the ones that would fit into RAM."

The boot loader starts in FLASH (obviously) and then loads a RAM-located portion into RAM which is used for actual FLASH programming. The initial FLASH code does a whole load of initialisation, verification, etc, not least because the RAM based portion has to get the code to program from "somewhere", in this case from a 4MB serial FLASH chip.

The above bit of code which got replaced with memcpy() was in FLASH. I had already spent many hours making sure that none of the C code is making calls anywhere outside the 32k boot block; that is a requirement if the boot loader is to be always capable of recovering a bricked device. I achieved this, but didn't bank on the bloody compiler doing the above stunt :) The basic issue here is that there is no way to defend against this "optimisation". I am fairly sure -Og doesn't do it (I checked this exactl piece of program, and the loop does remain. Will re-check with -O2, but frankly -Og is good enough and delivers a big size reduction (1/3) and looking at the assembler an equivalent speed-up too.

Good point re checking the map file for text objects outside the boot loader. It appears a little involved when the boot loader is made up of several .c files, and being a part of a bigger program memcpy() will exist legitimately. I would need a means of checking whether any code in the bottom 32k references any code outside that. I think your script is processing the .o files (the assembler listing) since the memcpy etc would not be visible elsewhere.

"A bootloader is a self contained project"

It usually is, apparently, but as you say I am unconventional :) I wanted to build this whole thing as one Cube IDE project, and it is pretty easy to do, and much easier to maintain long-term.

I now thing the best way is to develop in -Og and then switch to -O0 if you want meaningful single stepping to debug some algorithm. Otherwise the ITM data console is quite good for debugs (it crashes the STLINK debugger quite often though). Probably a dumb polled UART debug, 115kbaud, is the most reliable way.

EDIT: -O2 isn't much better than -O3. I am staying with -Og :)
Title: Re: GCC compiler optimisation
Post by: abyrvalg on August 16, 2021, 09:17:47 pm
I wanted to build this whole thing as one Cube IDE project, and it is pretty easy to do, and much easier to maintain long-term.
In the time spent figuring out all the necessary tricks for this strange requirement you could have polished a "conventional" bootloader into a rock solid code. But instead you have a code that still "echoes" at you (like that sudden memcpy in a wrong flash section), and that will continue.
harerod have suggested a way to create two binaries in one Cube project (by creating two build configurations with different defines/file sets/.ld and building both of them) more that a month ago IIRC.
Title: Re: GCC compiler optimisation
Post by: Nominal Animal on August 16, 2021, 09:31:30 pm
"Yes, because a compiler is allowed to turn code into a (standard) function call, when that standard function call is available, and -O3 enables all sorts of "unsafe" optimizations and assumptions the compiler may do to get stuff "run faster".  Usually it fails at it, though, so -O3 is rarely used."

That isn't exactly what those above have been telling me (that any failure due to optimisation means my code is crap) :) I can see this "optimisation" is legitimate, but that's not the same Q.
There is optimization, and then there is unsafe optimization.
-Os optimizes for size.  Often the code is fast as well.
-Og optimizes for debugging.
-O and -O1 enables optimizations.
-O2 optimizes even more.  This is the setting that vast majority of projects use, and some people mean when they tell you to compile with optimizations enabled.
-O3 optimizes yet more.  These optimizations usually cause code bloat, and often includes optimization features that have been relatively recently implemented, and are still being tested.  If they were always useful, they'd be included in -O2.  It should not enable any features that relax strict standards conformance, so if the compiler developers were perfect programmers, it would be safe to use -O3.  Unfortunately, in reality, -O3 tends to enable features that programmer-users and compiler-programmers disagree wrt. the standard, or are not sufficiently integrated or debugged in the compiler.  So, while -O3 is safe in theory, it is unsafe in reality.
-Ofast enables all (-O3) optimizations, plus some that are not strictly standards compliant.  That makes it unsafe.

When I advise new programmers, I always recommend starting with -Wall -O2.  Unfortunately, I too may occasionally say just "enable warnings and optimizations when compiling", when I actually mean -Wall -O2 exactly.  I also occasionally say "enable warnings and optimize for size when compiling" (especially for AVRs and other MCUs with very little RAM/ROM/Flash), when I actually mean using -Wall -Os.  I do apologize for this (and the others should too); it's just that these are so common, that anything else feels like an exception that needs describing.

If a compilation command includes more than one -O option, GCC uses the last one only.  Not all C compilers support -Og or -Ofast, or have similar descriptions for the different optimization levels; but in practice, -Os and -O2 tend to be the ones actually used nevertheless.

I think your script is processing the .o files (the assembler listing) since the memcpy etc would not be visible elsewhere.
Yes; it assumes that the Flash updating code is compiled in separate .c file or files, and compiled to a single .o object file, before being linked into a single ELF file that is then converted to hex for uploading to the device.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 16, 2021, 09:53:50 pm
and it is pretty easy to do, and much easier to maintain long-term.
It is not easier. Clear separation has concrete advantages. But it is too much work to convince people on the internet to do the right thing.

All your issues stem from that unconventional approach. You are not unique, many people have tried this. People simply rethink things more often when they see the downsides of a new and creative thing they are trying.
Title: Re: GCC compiler optimisation
Post by: DavidAlfa on August 16, 2021, 10:17:47 pm
Are you doing the whole thing in a single project? I doubt it's the way.
You have to make first the bootloader, setting the flash layout in the linker script, telling the compiler "you can only use these 12KB".
Then, a second program for the real application.
You have an option in the linker settings to generate relocatable code, so it'll use offsets instead literal addresses.
Now you could write that code anywhere in the flash and should work, at least in theory. Not sure about the vector table and such...

You don't need running in RAM. You erase the flash in pages, just don't wipe the bootloader's.
Although you would need a secondary bootloader to update the first uploader...

Also, you can make flash sections in the linker script, and tell the compiler where you want to place every function.
I don't know if it would still optimize, jumping to other flash addresses like the memcpy you saw.
Title: Re: GCC compiler optimisation
Post by: Doc Daneeka on August 17, 2021, 02:14:35 am
Quote
As to volatile, it's - unfortunately - more subtle than what you said.
1. To begin with: you talk about "observable". The only sentence actually using this term (in C99 at least) in the std is in this part:

An object that is accessed through a restrict-qualified pointer has a special association
with that pointer. This association, defined in 6.7.3.1 below, requires that all accesses to
that object use, directly or indirectly, the value of that particular pointer.
The intended
use  of  the restrict qualifier  (like the register storage  class)  is  to  promote
optimization, and deleting all instances of the qualifier from all preprocessing translation
units  composing  a  conforming  program  does  not  change  its  meaning  (i.e.,  observable
behavior
).

C11 onward explicitly adds that access to volatile objects is part of the observable behavior - so this is clearly what is intended here in C99 even though it is not spelled out - otherwise some core semantics of C would be changing between the standards - unlikely.

Quote
From what I understand here, this defines an "observable behavior" as the *meaning* of a program. Problem here is: what is the meaning of a program? The way I get this is the same as what I meant by the "functional POV", so anything volatile-related, when it may have unknown side-effects, but no analyzable effect, is NOT observable behavior. I may be wrong here and I admit we are really nitpicking on terms. I could not find the definition of the "meaning of a program" in the std.

As I said, C11 explicitly adds that access to volatile objects is observable behavior - assuming that is the intent in C99, any volatile access - even one with no side effects 'within' the C program, is observable.

Quote
For instance, taking the typical "delay loop" example, is a "delay loop", doing absolutely nothing apart from taking CPU cycles, part of the meaning of the program? If you can answer this one by a resounding "yes", without a blink, and backing it up with solid arguments, you are better than I am.

Without a definition of 'meaning' of a program - who knows? There are no CPU cycles in abstract C -  But one thing is certain it is not observable behavior as far as abstract C is concerned (assuming something like an empty loop with no volatile or library calls etc etc.). There is probably a reason the standard is framed in terms of observable behavior and not 'meaning'.

Quote
2. More importantly, about the volatile qualifier: it's unfortunately more subtle than it looks. Let's again quote C99 for the relevant parts:

Quote
An  object  that  has  volatile-qualified  type  may  be  modified  in  ways  unknown  to  the
implementation or have other unknown side effects.  Therefore any expression referring
to such an object shall be evaluated strictly according to the rules of the abstract machine,
as described in 5.1.2.3.

So far so good. Looks like "volatile" will guarantee that such a qualified object is evaluated in all cases, right?
But we need to refer to the "rules of the abstract machine" it mentions. So, again, relevant parts:

Quote
Accessing a volatile object, modifying an object, modifying a file, or calling a function
that does any of those operations are all side effects,
which are changes in the state of
the  execution  environment.  Evaluation  of  an  expression  may  produce  side  effects.  At
certain specified points in the execution sequence called sequence points, all side effects
of previous evaluations shall be complete and no side effects of subsequent evaluations
shall have taken place. (A summary of the sequence points is given in annex C.)

Still looks, at this point, like the volatile object will be evaluated no matter what. But the following paragraph kind of ruins it all:

Quote
In the abstract machine, all expressions are evaluated as specified by the semantics. An
actual implementation  need  not  evaluate part of an expression if it can deduce that its
value is not used and that no needed side effects are produced (including any caused by
calling a function or accessing a volatile object)
.

So, it looks a bit like what I said earlier. Doesn't it? (See the part in bold.)

No it doesn't - that paragraph does not ruin it at all - it says if something is not a needed side effect (which I don't think is defined anywhere, but it does not matter) it does not need to be evaluated - it absolutely does not say it must *not* be evaluated - you can still have other conditions on an expression which requrie that it be evaluated - which is what all the other paragraphs above do

- "An implementation does not need to evaluate every expression"
- "An implementation must evaluate an expression referring to a volatile qualified object"

It's pretty clear

In any case scope is irrelevant - a 'local' volatile variable for example might be implemented 'on the stack' but might be modified asynchronously - maybe another thread or an OS or task manager or something goes in and fiddles with it - there is even a footnote:

Quote
A volatile declaration  may  be  used  to  describe  an  object  corresponding  to  a  memory-mapped
input/output  port  or  an  object  accessed  by  an  asynchronously  interrupting  function. Actions  on
objects  so  declared  shall  not  be  ‘‘optimized  out’’ by an implementation  or  reordered  except  as
permitted by the rules for evaluating expressions.

The standard (the definition of the language - at least in modern terms) - cannot have anything to say about where local scoped variable are kept - it does not matter to the semantics all that matters is the C program cannot access them outside of their scope - inside their scope all the other semantics still apply
Title: Re: GCC compiler optimisation
Post by: ejeffrey on August 17, 2021, 02:15:53 am
"Yes, because a compiler is allowed to turn code into a (standard) function call, when that standard function call is available, and -O3 enables all sorts of "unsafe" optimizations and assumptions the compiler may do to get stuff "run faster".  Usually it fails at it, though, so -O3 is rarely used."

That isn't exactly what those above have been telling me (that any failure due to optimisation means my code is crap) :) I can see this "optimisation" is legitimate, but that's not the same Q.
There is optimization, and then there is unsafe optimization.
-Os optimizes for size.  Often the code is fast as well.
-Og optimizes for debugging.
-O and -O1 enables optimizations.
-O2 optimizes even more.  This is the setting that vast majority of projects use, and some people mean when they tell you to compile with optimizations enabled.
-O3 optimizes yet more.  These optimizations usually cause code bloat, and often includes optimization features that have been relatively recently implemented, and are still being tested.  If they were always useful, they'd be included in -O2.  It should not enable any features that relax strict standards conformance, so if the compiler developers were perfect programmers, it would be safe to use -O3.  Unfortunately, in reality, -O3 tends to enable features that programmer-users and compiler-programmers disagree wrt. the standard, or are not sufficiently integrated or debugged in the compiler.  So, while -O3 is safe in theory, it is unsafe in reality.
-Ofast enables all (-O3) optimizations, plus some that are not strictly standards compliant.  That makes it unsafe.

When I advise new programmers, I always recommend starting with -Wall -O2.

I would consider this advice on -O3 obsolete when targeting x86_64 and probably Armv8-A. On these type of architectures -O3 can show considerable performance improvements and the bugs and standard interpretation disagreements have been mostly worked out.  On a microcontroller it is probably true that -O3 does more harm than good.  On a general purpose CPU you should probably try both if you care a lot about performance.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 17, 2021, 06:17:08 am
Re the questions on what I am doing, see these:
https://www.eevblog.com/forum/microcontrollers/how-to-create-elf-file-which-contains-the-normal-prog-plus-a-relocatable-block/ (https://www.eevblog.com/forum/microcontrollers/how-to-create-elf-file-which-contains-the-normal-prog-plus-a-relocatable-block/)
https://www.eevblog.com/forum/microcontrollers/elf-to-binary-for-boot-loader/ (https://www.eevblog.com/forum/microcontrollers/elf-to-binary-for-boot-loader/)
https://www.eevblog.com/forum/microcontrollers/32f417-best-way-to-program-the-flash-from-ram-based-code/ (https://www.eevblog.com/forum/microcontrollers/32f417-best-way-to-program-the-flash-from-ram-based-code/)

-O0 produces 230k
-Og produces 160k
-O2 produces 160k
-O3 produces 180k
-Os produces 146k

I will try to incorporate Nominal Animal's script in the post build batch file. Can anyone recommend a good set of win32 command line utils (sed awk grep etc)? I have an old win16 set but they obviously don't run anymore. Even just getting a list of functions called from the boot block .c files would be sufficient; stuff like memcpy would be obvious. I already got caught with printf debugs but that should have been obvious, so wrote my own puts() :)

Incidentally, using the auto-CS on SPI feature would have saved a whole load of trouble, especially as my SPI2 is dedicated to the serial FLASH. And even with multiple SPI device on one SPI channel, using auto-CS and a demux driven from another couple of pins would work for multiple devices.
Title: Re: GCC compiler optimisation
Post by: DavidAlfa on August 17, 2021, 07:07:35 am
Maybe cygwin/mingw?
I use cygwin, comes really handy, but will never be the same as a Linux box.
For these commands, bash scripts, dd, sed... It's perfect.
Title: Re: GCC compiler optimisation
Post by: harerod on August 17, 2021, 04:36:54 pm
...
There is optimization, and then there is unsafe optimization.
-Os optimizes for size.  Often the code is fast as well.
-Og optimizes for debugging.
-O and -O1 enables optimizations.
-O2 optimizes even more.  This is the setting that vast majority of projects use, and some people mean when they tell you to compile with optimizations enabled.
-O3 optimizes yet more.  ...

When I advise new programmers, I always recommend starting with -Wall -O2.
...

Nominal Animal, thank you for this list. I have been a huge fan of "-Wall -Og", ever since the option became available. I prefer this setting as standard over the higher optimization levels, a way to make sure that the production code will fit the target. I only switch back to "-Wall -O0" when debugging gets nasty.
In your understanding - what would be the drawbacks of using -Og in production code? Size- and performance-wise I see not much difference to -O2.
Title: Re: GCC compiler optimisation
Post by: Nominal Animal on August 17, 2021, 08:47:18 pm
In your understanding - what would be the drawbacks of using -Og in production code? Size- and performance-wise I see not much difference to -O2.
For open source, none really.  For proprietary code, the impact of leaving the debugging information in the binaries is something to consider.  (Sometimes it is a positive, sometimes PHBs don't like the idea at all.  It varies.)

On embedded targets (appliances and microcontrollers) where size is an issue and you do not usually do any debugging on release hardware, optimizing for size and stripping unneeded symbols from ELF binaries (at release, or before converting to a hex file to be uploaded) may be the difference between things fitting and working, or not.  On such situations, I usually make it easy to rebuild the binaries, and test the compiler version used for the actual effects.  Once again, reality wins over theory.

If someone asked me why I am not recommending -Og instead, I'd have to admit, "old habits".  Clang supports -Og as well with similar semantics (enable basic optimizations but with the focus on debuggability), so perhaps -Wall -Og would be a superior suggestion to new programmers; I am not certain yet.



In Linux distributions, there are usually separate versions of library binaries that have debugging information enabled.  It is also possible to provide the debugging symbols for a dynamic library, for example the standard C library (libc6), in a separate ELF dynamic library that only contains the debugging information (libc6-dbg in Debian derivatives).

I don't usually need debugging information in my binaries.  This is not to say I don't do it, only that I have a lot of tools I can use instead, and gdb just isn't one of the fastest/most efficient ones for most cases for me.  When I do use e.g. gdb, I do like to make things easy for myself, for example via python pretty-printing gdb extensions (https://stackoverflow.com/a/23970415).

As an example, consider abstract data types, like trees, graphs, heaps and such.  When I implement a new (to me) one, I always create a test program that includes graph/tree/heap traversal that does not get confused by loops (i.e., stores the pointer of each visited node, or marks each node visited if there is room in the data structure), and have it emit a nice Graphviz Dot (http://graphviz.org/documentation/) format description of the data structure.  Dot is a very simple, but powerful and expressive text format, so very easy to emit from ones code.  I then test it with random, typical, and pathological data sets, and examine the trees such generated.  Not only is it obvious if there is an unwanted cycle or similar problem, but adding information (like recursion depth, level, or distance from initial node) to the Graphviz graph can help pinpoint the root cause in fraction of the time than e.g. single-stepping through the code could.

Even when writing say recursive code, emitting the call graph as a Graphviz Dot directed graph, can be much more informative than gdb debugging, even single-stepping through the recursive code.  See this (https://www.nominal-animal.net/answers/fibo.svg) and this (https://www.nominal-animal.net/answers/fibonacci-4.svg) for example graphs (in SVG form, drawn using graphviz dot -Tsvg then cleaned up in Inkscape and minimized by hand for online use), describing how the very common recursive Fibonacci exercise call graph can be visualized.  If you are familiar with the sequence and the exercise, I don't think I even need to describe the graphs, really; it is pretty darned obvious...  just imagine if you had a bug, how obvious that bug would be in the graph.

I believe the visual methods also helps new programmers to come up with their own visualization/modeling methods, when dealing with code or data structures and flow.  (My own tend to be chaotic, as if I were using paper as a cache for my mind.  Because of that, after I work a problem out, I write a simple text file, perhaps with a descriptive image or two, to document the solution, saving them and the code in a dedicated directory.  After a month or a year, the stuff is as foreign to me as it would be to anyone else.  I don't like trying to memorize anything, so that documentation is useful even if I were the only one ever to access them.)

All this said, I would not be surprised if someone had a different experience and therefore different opinion on this.  This is just one of the patterns I've found to work.
You could say a core reason why I place much less weight on debugging tools than others is that I very much believe understanding (or "grokking") the intent/purpose/design of the code correctly, is much more important than getting the code to have the effects you want.  Debugging tools help you with the latter, but I want to do the former, and emphasize the need to do the former, for example via visual tools like Graphviz.
Title: Re: GCC compiler optimisation
Post by: ttt on August 17, 2021, 09:13:56 pm
I take a slightly different perspective on optimizations with gcc, it's not all about -Ox. I am usually actively trying to fit the code into the lowest cost MCU possible, which can mean the smaller flash size variant.

So -Og for debugging and -Os for release builds by default. And then decorate my performance critical function with optimization attributes as such:

    __attribute__ ((hot, flatten, optimize("O3"), optimize("unroll-loops")))
    void Strip::ws2812_alike_convert(const size_t start, const size_t end) {
...

In addition, and that has not been mentioned in this thread yet, link time optimization (-flto) makes a _huge_ difference code size and performance wise. Though it can be tricky to make a code base LTO safe, link order and symbol visibility issues can be frustrating.
Title: Re: GCC compiler optimisation
Post by: cfbsoftware on August 17, 2021, 09:36:32 pm
So -Og for debugging and -Os for release builds by default.
Do you do your both your unit testing and integration testing on both builds?
Title: Re: GCC compiler optimisation
Post by: ttt on August 17, 2021, 10:04:45 pm
So -Og for debugging and -Os for release builds by default.
Do you do your both your unit testing and integration testing on both builds?

Yes, if it fits :-) Behavior of the code changes depending on optimization flags. You would think that it is a sign of badly written code but it is all to common to run into race conditions if anything is slightly timing critical.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 17, 2021, 11:58:44 pm
Can anyone explain what is meant by "debugging information" included in the code?

I am able to step through code compiled with -O0 all the way through the -O3 etc. With the higher levels it doesn't make much sense but the source code is still visible because it is available locally so the debugger can refer to it.

" it is all to common to run into race conditions if anything is slightly timing critical."

It isn't just race conditions; it is all kinds of stuff which break. In addition to the example already posted (replacement of a loop with memcpy() which broke a program which was supposed to live in the bottom 32k) I have just spent a few more hours chasing down a much more subtle issue where an Adesto serial FLASH seems to have an undocumented sensitivity to minimum CS=1 time when reading out parts of the manufacturer ID etc.

I have cygwin (use rsync a lot) so will try that for the above mentioned test.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 18, 2021, 12:03:28 am
There is no debugging information in the code itself. You should not confuse options -Og, which is just optimization setting that generates code that is better for debugging (less rearranging). And -g option, which generates actual debug information (correspondence of the assembly instructions to the C source, variable names and locations, etc).

In any case you can strip the debug information from any ELF file regardless of how it was created.

And if you only distribute binary files (BIN, HEX), then there is no debug information, of course. It only applies to ELF files.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 18, 2021, 12:10:08 am
Yes, debug information is generated with the '-g' option. That can be combined with any optimization options.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 18, 2021, 12:52:09 am
Having done 40 years of assembler I remain unimpressed with that loop replacement with memcpy.

If the loop is short then the replacement cannot make sense. The only thing which would help, and only with a prefetch queue, would be to unroll the loop and inline it. The memcpy function is a lot of code because it will do it 4 bytes at a time and then (or beforehand) tidy up any unaligned ends. And if the compiler was trying to evaluate the loop length, it got it wrong because it was at most 6 bytes (within a 512 byte buffer).

If the loop is long then ok. But the designer would have probably used memcpy anyway if doing hundreds of bytes or more.

I see FatFS do their own versions of these functions, probably because they had problems too. But the compiler will try replacing these too.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 18, 2021, 01:20:41 am
If the loop is short then the replacement cannot make sense.
But if you use that memcpy in many cases, then it will be replaced with calls. If you only have one, then it makes no sense to place it as a function and them make a call.

And doing aligned transfers is good for performance.

I see FatFS do their own versions of these functions, probably because they had problems too. But the compiler will try replacing these too.
Their functions would be recognized as copy functions and replaced with memcpy(). They did it because you can't guarantee that memcpy() will be present on the target platform. I often build things with "-nostdlib" flag, making sure no standard library functions are included.

Title: Re: GCC compiler optimisation
Post by: newbrain on August 18, 2021, 09:06:36 am
Slightly off-topic, as we are talking about gcc here, but here is an interesting explanation on how clang + LLVM perform optimizations (https://blog.matthieud.me/2020/exploring-clang-llvm-optimization-on-programming-horror/).

The example is this nifty test for evenness:
Code: [Select]
bool isEven(int number)
{
    int numberCompare = 0;
    bool even = true;

    while (number != numberCompare)
    {
        even = !even;
        numberCompare++;
    }
    return even;
}

Do not miss the corresponding Hacker News discussion (https://news.ycombinator.com/item?id=28207207).

The takeaway is, again: a conforming C compiler is allowed to do more or less whatever it fancies, as long as the osservable behaviour (in this case: calling the isEven function, under a number of assumption...) is guaranteed.
Title: Re: GCC compiler optimisation
Post by: harerod on August 18, 2021, 09:39:48 am
Quote from: harerod on Yesterday at 17:36:54

    In your understanding - what would be the drawbacks of using -Og in production code? Size- and performance-wise I see not much difference to -O2.

...
If someone asked me why I am not recommending -Og instead, I'd have to admit, "old habits".  Clang supports -Og as well with similar semantics (enable basic optimizations but with the focus on debuggability), so perhaps -Wall -Og would be a superior suggestion to new programmers; I am not certain yet.
...

Nominal Animal, I appreciate your input, because of our different views. You seem to be a highly trained software expert who happens to write code for embedded systems. I am a hardware designer who also happens to write code. :)

+ + +

Several posts ago somebody asked for a tutorial. Why not have a look at the available documentation? Basic information for this thread spans STM32, CubeIDE and ARM-GCC:

First of all the MCU involved:
https://www.st.com/en/microcontrollers-microprocessors/stm32f407-417.html#documentation (https://www.st.com/en/microcontrollers-microprocessors/stm32f407-417.html#documentation)

Maybe the release note for the CubeIDE:
https://www.st.com/content/ccc/resource/technical/document/release_note/group0/9a/72/48/16/ec/bd/44/5a/DM00603738/files/DM00603738.pdf/jcr:content/translations/en.DM00603738.pdf (https://www.st.com/content/ccc/resource/technical/document/release_note/group0/9a/72/48/16/ec/bd/44/5a/DM00603738/files/DM00603738.pdf/jcr:content/translations/en.DM00603738.pdf) <- RN0114 CubeIDE release note

ARM-GCC:
https://developer.arm.com/tools-and-software/open-source-software/developer-tools/gnu-toolchain/gnu-rm (https://developer.arm.com/tools-and-software/open-source-software/developer-tools/gnu-toolchain/gnu-rm)
https://gcc.gnu.org/onlinedocs/9.3.0/ (https://gcc.gnu.org/onlinedocs/9.3.0/) <- GCC used in CubeIDE 1.7.0, if I read RN0114 correctly
Title: Re: GCC compiler optimisation
Post by: DavidAlfa on August 18, 2021, 10:50:20 am
Not that kind documentation. A proper guide/manual for the HAL.
You have a simple one, barely describing what each function does. But you still have to guess a lot of things.
Title: Re: GCC compiler optimisation
Post by: harerod on August 18, 2021, 11:21:13 am
DavidAlfa, let's drift too far off-topic. Kindly follow me to the CubeIDE thread:
https://www.eevblog.com/forum/microcontrollers/is-st-cube-ide-a-piece-of-buggy-crap/msg3632875/#msg3632875 (https://www.eevblog.com/forum/microcontrollers/is-st-cube-ide-a-piece-of-buggy-crap/msg3632875/#msg3632875)


Title: Re: GCC compiler optimisation
Post by: Nominal Animal on August 18, 2021, 02:22:02 pm
I would consider this advice on -O3 obsolete when targeting x86_64 and probably Armv8-A.
Perhaps, but it very much depends on the compiler and especially compiler version.

For example, I do not use anything newer than GCC 9 for ARM targets, because of the unfixed issues in later versions.  I'm seriously considering switching to Clang for arm, anyway.

I take a slightly different perspective on optimizations with gcc, it's not all about -Ox.
Very true; I too mentioned specific optimization flags that I end up using; for example, -ffinite-math-only can make a big difference and be very useful when you have e.g. explicit checks so that you never do division by values very close to zero and such.

However, I like to keep such things in separate compilation units (files).  For optimized routines, I often have alternates with the exact same interface but wildly different implementations, and choose the implementation simply by selecting which C source file (among the alternates) is used: either via Makefile options, or via a common .c source file that #includes the appropriate .c source file based on preprocessor macros.  (Note that some people do have an irrational dislike of #include used with source files though; it seems that it jars some peoples sensitivities somehow.)

Can anyone explain what is meant by "debugging information" included in the code?
As ataradov and SiliconWizard already mentioned, there is no debugging information per se in the code.

Some optimization flags do affect the debuggability of the code, though; in particular, -fomit-frame-pointer.  In many architectures, the address of the current stack frame is kept in a separate register.  This option disables that (so that the stack frame is then implicit, and local variables on stack are accessed via the stack pointer).  On some architectures, this can make debugging much harder; according to documentation, impossible on some, but I'm not sure on which arches that is.  Stack frames can still be described for each function via separate debugging data, for example when using DWARF formats (for the debugging data).

Object files and especially final ELF binaries will contain a lot of extra information when debugging information is enabled.  If your build facilities are such that the ELF files are stripped before uploaded to the target device (say, like in Arduino environment, or most environments targetting microcontrollers), it does not matter whether the compilation included debugging information in the object files or not (whether -g was used or not); but the optimization options used does.

One nice thing about -Og is that the compiled code should be the same regardless of whether debugging information is included or not via -g, in the object files and final binaries.  If one uses -O2 or -Os , and then switches to say -O0 -g for debugging, the compiled code is usually different, making debugging problems more difficult than necessary.  Also, both -O2 and -Os enable -fomit-frame-pointer, affecting debuggability.  I do believe it was the programmers' need for an optimization level that generates reasonably optimized code without affecting debuggability that caused GCC to grow support for -Og in GCC 4.8 in 2013, I believe; but I'm not sure if Clang actually implemented it first and GCC users found how useful it is, or vice versa.

Since we're using the ELF file format for object files (and final binaries before converting to hex), knowing the structure of ELF files (https://en.wikipedia.org/wiki/ELF_file_format) can be very useful, since ELF files can contain all sorts of information, not just "code" and "data".

I often build things with "-nostdlib" flag, making sure no standard library functions are included.
Me too, but with GCC, it is not enough to avoid a dependency on memcpy(), memmove(), memset(), and memcmp(), because GCC expects these to be provided by even a freestanding environment; see the second-to-last paragraph in section 2.1, GCC C-Language Standards (https://gcc.gnu.org/onlinedocs/gcc/Standards.html#C-Language) in GCC documentation.

It is not too common for GCC (across its versions and compiler options) to turn loops into a call of one of the above, I think.

A bit of glue logic (even preprocessor macros detecting compile-time type or alignment, so that an optimized native-word-sized operations can be used) is usually enough to ensure it does not happen for a particular function implementation.  When one does need these four functions anyway, or duplicates of them in the same project, the library-provided ones are weak, and one can override those simply by implementing ones own, using the same function signature (including name).

If one defines them in the same compilation unit (file or files compiled in the same gcc/clang command), the pattern shown by ttt in post #206 (https://www.eevblog.com/forum/microcontrollers/gcc-compiler-optimisation/msg3632285/#msg3632285) can be used to control the optimization flags; and the scriptlet I showed earlier can be used to verify the compiled object file contains no external dependencies.

Sometimes it can be worth the effort to implement these (separate variants for loop direction and access size for memcpy()/memmove(), separate access sizes for memset(), and only a byte-by-byte memcmp()) in extended inline assembly (https://gcc.gnu.org/onlinedocs/gcc/Extended-Asm.html#Extended-Asm) (asm volatile ("code" : outputs : inputs : clobbers); as the function body).  If in the same compilation unit, I recommend using an #include "memfuncs.c" so that the implementation is easy to change/select at build time, for example based on the hardware architecture.  It also makes unit testing them (with a separate program) much easier.

Nominal Animal, I appreciate your input, because of our different views.
I too appreciate different views, especially when people describe the reasons for their different views (like you did, and members like ataradov and SiliconWizard and many others do), because that way I can learn.  I know I don't know really that much; but I can and am willing to learn.  Never hesitate to correct me, if you believe I am in error; I very much appreciate that.

My communications style is far from optimal (verbose, sometimes looks like I'm trying to be more authoritative than I actually am, me occasionally fail English, and so on); but my attempt is always to describe my reasons, with my current opinion (based on those reasons) more like a side note than the focus, because my opinions change as I learn.  But that sort of describing-the-reasons can sometimes appear as The List Of Facts, which they aren't; they're just the stuff I currently am aware of.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 18, 2021, 04:58:27 pm
"My communications style is far from optimal "

Your communication style is excellent :)
Title: Re: GCC compiler optimisation
Post by: Nominal Animal on August 19, 2021, 06:55:24 pm
No.  If I were more concise and less blunt and confrontational, I'd be more effective.

As it is, I'm in the ignore list (Profile, Modify Profile > Buddies/Ignore list... > Edit ignore list) of quite a few members and at least one admin.

I've managed to anger several members whose knowledge and expertise I value, inadvertently, due to my communications style – ataradov and Bruce Hoult I know for sure, but how many others I haven't even realized? :-//

I should be more aware, because I use my pseudonym for the express purpose of being able to interact with others: I'm much too sensitive to perceived slights to my person, and using a pseudonym helps me remember any negativity is caused by my own output, and is not about my immutable characteristics: about what I said, not about who I am.  My output I can affect, at least to some degree; my person, not so much.



As to GCC compiler optimization:

GCC has version-specific documentation (https://gcc.gnu.org/onlinedocs/), as well as the latest version of the documentation (https://gcc.gnu.org/onlinedocs/gcc/) available online.  Don't be afraid or think you need to know these, because you don't.  You don't memorize IC datasheets, so why the heck would you memorize GCC or Make manuals either?  If you use them constantly, some details may stick in your memory, and that's fine; but it isn't necessary at all.  My own memory does not do that much (or rather, it can get the smallest details wrong), so I personally don't trust my memory for the details, and instead have concentrated on my searching and lookup skills (including fast reading/glancing to find appropriate contexts, that I need to actually read, in the sense that I am conscious of the sentences).  I usually have a browser window (in a separate workspace/virtual desktop) with tabs open to relevant manuals only.  (I can warmly recommend the Linux man-pages project (https://www.kernel.org/doc/man-pages/) for up-to-date pages on POSIX C interfaces.  It is not complete wrt. non-Linux interfaces (BSD, Solaris), but those that it does cover, it mentions even the standards/sources where those interfaces are derived.)

In this thread, the GCC documentation page on optimize options (https://gcc.gnu.org/onlinedocs/gcc/Optimize-Options.html#Optimize-Options) should be extremely useful.  Not just the basic -O options, but also wrt. the individual options one might define for specific files or functions.

I started with systems integration stuff ("making my own Linux distro") before the turn of the century, in the Linux from Scratch (https://linuxfromscratch.org/) community.  Around the turn of the century, I also started packaging some of my output as RPM and DEB packages.  Even if one does not intend to create such packages, browsing through the Debian packaging tutorial (https://www.debian.org/doc/devel-manuals#packaging-tutorial) (and say RPM packaging guide (https://rpm-packaging-guide.github.io/)) is an excellent source of an overview of how software packages have been delivered on a number of Linux systems very efficiently.  Aside from human causes (incorrect/bad package dependencies and such), these are very robust formats, and include pre-install, post-install, pre-remove, and post-remove shell scripts triggered by the corresponding action: this is where I discovered the utility of such, and started incorporating scripts into my Makefiles, so that I could more easily fully automate my builds.

Now, if one goes to some package page in Debian or debian-derivatives like Ubuntu, say Inscape for Ubuntu 20.04 LTS (focal) (https://packages.ubuntu.com/focal/inkscape), you can find the source archives for the package.  If you download the .debian.gz/.debian.bz2/.debian.xz one, you get the Debian packaging additions (that are extracted into the debian/ subdirectory of the original source tree).  The interesting file in source archives is the rules file, since it defines the variables (DEB_BUILD_MAINT_OPTIONS, DEB_CFLAGS_MAINT_APPEND, DEB_LDFLAGS_MAINT_APPEND) describing the options (including optimization options) on how the sources were compiled to obtain the particular binaries.  If not defined, the defaults set in the project source tree (Makefile, CMake, etc.) are used.  Similarly, the spec file in source RPM files (.srpm) contain the optimization options etc. used to compile the binaries; see e.g. Fedora build flags (https://src.fedoraproject.org/rpms/redhat-rpm-config//blob/rawhide/f/buildflags.md) documentation for details.

There are three methods to modify the optimization options within a single source file: the optimize function attribute (https://gcc.gnu.org/onlinedocs/gcc/Common-Function-Attributes.html#index-optimize-function-attribute), #pragma (https://gcc.gnu.org/onlinedocs/gcc/Function-Specific-Option-Pragmas.html), and _Pragma() (https://gcc.gnu.org/onlinedocs/cpp/Pragmas.html).  GCC documentation states that these may not support all options, and should only be used for debugging, not production code.  Because of this, I do recommend compiling functions that may need special compile options in separate compilation units (separate .c source files) instead.

Since makefiles support both generic recipes and recipes used for specific files (even though they match the generic recipe rule), one only needs to add a new compilation recipe to handle each specific file or files that need separate compiler options to be used.

The GNU make manual (https://www.gnu.org/software/make/manual/) is something any Makefile user should at least browse through.  Not to memorize anything, or even understand the purpose, but to get an overview of its capabilities.  Like with source control tools, you don't need to be a PhD to use make effectively; the important thing is to understand the overview, and have the manual handy for referencing any details.

The linker and linker scripts are another thing that initially seem very complex, but are actually rather straightforward.  The terminology (section vs. segment, and so on) used in ELF linkers can be confusing, so I recommend writing a short crib sheet for the terms, using ones own words.  With GCC and Clang, I do not recommend directly executing the linker at the link phase; it is better to let the compiler internally call the (correct) linker for the target.  Both compilers do provide command-line options on how to pass parameters to the linker (-Wl,param as a single command-line argument, so this is not a restriction, really.  This also means that if a Makefile contains $(LD), you know it is an "old-style" one that executes the linker directly, instead of through the compiler.  In particular, the compiler knows the target and options needed to supply to the linker to generate code for that target, so there may be some parameters not passed to the linker if you execute the linker directly, or you might even execute the incorrect linker.  Better let the compiler handle it.
Title: Re: GCC compiler optimisation
Post by: Siwastaja on August 20, 2021, 06:46:49 am
Nominal, this is OT but just a quick note, when people discuss, sometimes things get a bit heated and people even get angry. This usually means no hard feelings afterwards, and when kept to modest amounts, is felt rewarding or cathartic. It's normal, it's life, and something which makes life more worth living, I would not prefer a dull world where emotions are all but suppressed. I'm 99% sure you haven't angered ataradov or Bruce Hoult "in the wrong way" at all.

Those who can't grok that you also have an emotional side as well, it's their problem. If you are on their ignore lists, good riddance, at least they won't pick unnecessary battles.

Not all long posts are created the same. Your verbose style is OK because I find it easy to just skim through the posts and still find the key points whenever I don't have time to read it all. This is because despite verbosity (and occasional addition of OT), the posts are properly organized and paragraphed. And when I do have time, I may enjoy reading it all slowly. The choice is left to the reader. No issue here.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 22, 2021, 08:52:42 pm
Back on the topic of compiler optimisation:

I've just watched this video:

https://www.youtube.com/watch?v=5GhHpeRgceM&ab_channel=ViktorVano (https://www.youtube.com/watch?v=5GhHpeRgceM&ab_channel=ViktorVano)

He seems to have used "volatile" for just about every variable in any way connected with the flash programming code, and I just don't get it. Take a look around 8:20.

(https://peter-ftp.co.uk/screenshots/202108221812904321.jpg)

Title: Re: GCC compiler optimisation
Post by: ataradov on August 22, 2021, 08:59:25 pm
This is absolutely not necessary, but also does not hurt. He still uses high level APIs, so all those volatiles will be discarded anyway.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 22, 2021, 09:06:19 pm
Whenever I tried declaring say a buffer as "volatile", and passed its address to some function whose prototype didn't have a "volatile" on that parameter, I got a compiler warning that the "volatile" is being discarded.

Sure it does not hurt but then why not just make everything "volatile". It makes no sense to do it as a precaution, just in case. AIUI, if you read location 0x0800000 into some variable then that variable must be "volatile", but code which subsequently accesses that variable does not have to be.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 22, 2021, 09:41:28 pm
Well, I did not say that you should do that, I just said that it does not hurt in this case. Of course correctly written code would not use a single volatile here. And yes, compiler would complain, and generally it is a very good idea to listen to those complaints.

If you read something into a variable, this variable does not need to be volatile. This is a perfectly valid code:
Code: [Select]
uint32_t variable = *((volatile uint32_t *)0x0800000);
Title: Re: GCC compiler optimisation
Post by: peter-h on August 22, 2021, 10:11:48 pm
OK; you declared the source as volatile. But nothing else should need it, so long as the origin of the data is declared thus.

How would you totally avoid "volatile" in this scenario?
Title: Re: GCC compiler optimisation
Post by: ataradov on August 22, 2021, 10:18:45 pm
You can't. The source must be volatile.
Title: Re: GCC compiler optimisation
Post by: Nominal Animal on August 22, 2021, 10:31:31 pm
Sure it does not hurt but then why not just make everything "volatile".
You could, but then the compiler could not optimize the code much.  The code would work as intended, but be slower than necessary.

It makes no sense to do it as a precaution, just in case.
No, but a lot of "programmers" throw stuff at the wall, and see what sticks.  The honest ones will tell you that "this is how I got it to (seem to) work, but I don't know how or why".  I consider them similar to software engineers as alchemists are to chemists.

AIUI, if you read location 0x0800000 into some variable then that variable must be "volatile", but code which subsequently accesses that variable does not have to be.
I'd prefer putting it a bit different way (since I'm not sure if that sentence describes the situation correctly or not; my English fails me here).

You declare a variable (or an object like an array) volatile, when you do not want the compiler to infer its value in any way, and want the compiler to access the actual storage (memory) of it whenever the variable or object is accessed.

If you look at ataradovs above example, you'll see the pattern of how to make a single access volatile: you construct a pointer to the volatile data, with the pointer pointing to the object or address you want, and then dereference the pointer.  The value is stored in a non-volatile variable.  To aid us human programmers, I like to explicitly declare such "locally cached values" as const –– which is just a promise to the compiler that we only read the value, and will not try to modify it, making it easier for the compiler to optimize the code using the const values.  (GCC and Clang are pretty darned good at inferring constness on their own, though, so in practice, it really is more a reminder to us humans that this value will stay constant in this scope.)

volatile is not the only way to tell the compiler that its assumptions about variables or objects are no longer valid.  In GCC, on all architectures, asm volatile ("" : : : "memory"); acts as a compiler memory barrier (generates no machine code itself!) that tells the compiler that any assumptions about memory contents across that statement are invalid.  However, it does not affect local variables and objects, since these are on stack.

Another way is to call a function whose prototype is known but implementation unknown – for example, compiled in a separate unit (C source file) –, that takes a pointer to the non-const memory range containing one or more variables.   For a specific variable or object (including a dereferenced pointer), you can also use asm volatile ("" : "+m" (object)); which tells the compiler that any assumptions about the value of object become invalid at that point.

How would you totally avoid "volatile" in this scenario?
For example,
        asm volatile ("" : "+m" (*(uint32_t *)0x0800000));
        uint32_t  variable =     *(uint32_t *)0x0800000;
        asm volatile ("" : "+m" (*(uint32_t *)0x0800000));
does the exact same thing.

In pure C, without inline assembly, there is no exactly equivalent alternate to volatile whose behaviour the standard guarantees.  In practice, bracketing the access with a call to a function, say foo((uint32_t *)0x0800000); with the implementation of foo() not visible to the compiler (but prototype e.g. void foo(uint32_t *); , the compiler would have to load the 32-bit unsigned integer at that point in the code.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 23, 2021, 06:26:31 am
Doesn't declaring a variable as extern also do it? The compiler cannot "see" across multiple source files.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 23, 2021, 06:35:10 am
Not necessarily, especially if you use LTO, which is a good idea.

Also, even example foo((uint32_t *)0x0800000);  is not entirely correct. If foo() writes to the flash, then sequential call to foo(); may use old value cached instead of reading the new one again. Or if something else writes to the flash and you want foo() to notice the change.

There is no need to invent new ways to trick the compiler. There are well defined ways to communicate what you want.
Title: Re: GCC compiler optimisation
Post by: Nominal Animal on August 23, 2021, 04:48:20 pm
If foo() writes to the flash, then sequential call to foo(); may use old value cached instead of reading the new one again. Or if something else writes to the flash and you want foo() to notice the change.
True.

There is no need to invent new ways to trick the compiler. There are well defined ways to communicate what you want.
Absolutely.  I use volatile rather often (in POSIX userspace applications, the most typical one is an interrupt flag, a volatile sig_atomic_t flag), and compiler memory barriers and such constructs extremely rarely; only when doing odd stuff with memory caching controls and such.

As a mental model for a human programmer, however, I do believe my examples of "compiler cannot assume value if ..." are useful, though.  (That is the way and reason I included them, at least; not as a practical example.)

They explain why you do not usually have volatile pointers in function prototypes (except static inline accessor functions compiled in the same unit/source file –– for GCC and Clang, these are just as fast as macros, with their function bodies included in the call site, but unlike macros, also provide type checking at compile time), as it'd really only affect the function implementation, for example.  They should also help see one way how C compilers optimize various expressions, by tracking the access to variables and avoiding superfluous accesses; and what that means to actual code generated, and how to avoid any problems from such optimizations.
(Which, if I've understood correctly, is the entire reason for this thread: having C code behave in unexpected/unintended ways when compiling with optimizations enabled, and trying to understand why and how that happens.)
Title: Re: GCC compiler optimisation
Post by: peter-h on August 24, 2021, 10:00:48 am
OK here goes another dumb Q:

(https://peter-ftp.co.uk/screenshots/202108243112925410.jpg)

The two g_ flags are picked up by another RTOS task. They are not referenced in the current .c file.

Should they be "volatile"? They are not declared as extern in the current .c file (but are in the other one, obviously) but are set to zero when defined.

Amazingly this code runs as-is even with -Og, and the compiler is not complaining about an unused variable.
Title: Re: GCC compiler optimisation
Post by: harerod on August 24, 2021, 12:11:11 pm
Have you considered using inter-task communication mechanisms, to make results less volatile?

https://freertos.org/Embedded-RTOS-Binary-Semaphores.html (https://freertos.org/Embedded-RTOS-Binary-Semaphores.html)
Title: Re: GCC compiler optimisation
Post by: peter-h on August 24, 2021, 12:58:40 pm
Yes but I believe in simplicity :)

I take it the answer to my question is Yes :) But maybe not?
Title: Re: GCC compiler optimisation
Post by: gf on August 24, 2021, 05:05:00 pm
Volatile can only provide limited memory ordering guarantees anyway. Memory barriers are IMO the better instrument to provide the memory ordering guarantees typically required in multi-threaded environments. In a single-processor environment, asm volatile("" ::: "memory") can act as compiler barrier, but in multi-processor environments even hardware memory barriers are frequently unavoidable (and volatile alone would no longer suffice then at all, even if all objects were volatile).

Various thread synchronization functions provided by the OS happen to be implicit memory barriers. For instance, if you protect the access to shared objects (i.e. shared between threads) with mutexes, then you get implied memory barriers at the points where you acquire and release the mutex. So if you use the OS-provided thread synchronization primitives, then you don't need to care about the low-level stuff, and you can renounce explicit memory barries, atomic operations or volatile qualifiers for shared objects in most cases.

Title: Re: GCC compiler optimisation
Post by: ataradov on August 24, 2021, 05:14:47 pm
The two g_ flags are picked up by another RTOS task. They are not referenced in the current .c file.
Should they be "volatile"? They are not declared as extern in the current .c file (but are in the other one, obviously) but are set to zero when defined.

From the compiler point of view tasks are just functions, and as long as there is a write and a read, it can't be optimized. Volatile here is useless.

But you need to be very careful how you use shared variables like this. It is safe to use simple types like booleans in some cases, but you really need to know what you are doing.

And if you start to get a lot of those shared variables, then you should be expecting for something to break.
Title: Re: GCC compiler optimisation
Post by: gf on August 24, 2021, 05:56:12 pm
The two g_ flags are picked up by another RTOS task. They are not referenced in the current .c file.
Should they be "volatile"? They are not declared as extern in the current .c file (but are in the other one, obviously) but are set to zero when defined.

From the compiler point of view tasks are just functions, and as long as there is a write and a read, it can't be optimized. Volatile here is useless.

The actual point in this code snipplet is whether the store to the g_ variables can be re-ordered behind the call to xTaskCrate(), or not.
If they are global varibles and the compiler cannot see what xTaskCrate() does, then it must assume anyway that they might be accessed in xTaskCrate(), therefore re-ordering behind the call can't happen.
But if the compiler can inline xTaskCrate(), or if LTO is used, then it depends...
[ Basically xTaskCrate() is one of these functions which IMO should act as implied memory barrier, as it spawns a new thread. ]
Title: Re: GCC compiler optimisation
Post by: peter-h on August 24, 2021, 06:22:07 pm
" It is safe to use simple types like booleans in some cases, but you really need to know what you are doing."

I am well aware that passing a unit64_t between RTOS tasks is possibly not safe, because it won't be written in one operation, and of course strings are worse. So if doing that, one needs an atomic flag to indicate when ready, etc.

I use the FreeRTOS mutexes around things like set/read RTC, because I am setting the RTC from the GPS in one thread and reading it from other(s). And one could do the same around simple variables. I've done it around all SPI3 ops because SPI3 is shared between different devices (and yes amazingly it does work, with some precautions).

I've measured the exec time of the mutex call and it is very fast - 1us IIRC.

AFAIK ARM 32F4 booleans are single byte, but even 32 bit vars would be atomically written, and read.
Title: Re: GCC compiler optimisation
Post by: peter-h on August 30, 2022, 06:54:25 am
I posted this in a thread which appears to have gone dead, and this is more on-topic.

Can I be sure that this loop will not be replaced with memcpy, due to the use of "volatile"?

(https://peter-ftp.co.uk/screenshots/202208295215661807.jpg)

Is isn't getting replaced (I am using -Og) so the answer is probably affirmative.

It actually could be a memcpy (because the CPU FLASH is not going to change) and I have a local version of memcpy which has optimisation turned off, and it would perhaps be better maintenance-wise to use that, but I am not sure of the syntax :)

addr starts at 0x8000000 - the base of FLASH.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 30, 2022, 07:04:46 am
It depends. With the current compiler -Og is close to -O1 with minor modifications. But this can change in the future.

The loop replacement is controlled by -ftree-loop-distribute-patterns flag, which is included by default only at -O3. EDIT: It looks like in recent versions it is also enabled in -O2. So, it already changed at least once.

But you can just be explicit and pass "-fno-tree-loop-distribute-patterns" to the compiler and it will not do the replacement at any optimization level.
Title: Re: GCC compiler optimisation
Post by: westfw on August 30, 2022, 07:45:53 am
Quote
you can just be explicit and pass "-fno-tree-loop-distribute-patterns" to the compiler
Or stick it in a pragma, for just that function?
Title: Re: GCC compiler optimisation
Post by: peter-h on August 30, 2022, 07:49:38 am
Quote
The loop replacement is controlled by -ftree-loop-distribute-patterns flag, which is included by default only at -O3. EDIT: It looks like in recent versions it is also enabled in -O2. So, it already changed at least once.

I have found memcpy replacement with -Og, IIRC. I experimented with levels other than -Og only briefly. We had a thread about it (maybe this one) way back. -O3 caused some problems, for no benefit.

Quote
But you can just be explicit and pass "-fno-tree-loop-distribute-patterns" to the compiler and it will not do the replacement at any optimization level.

I will do that - thanks.

Quote
Or stick it in a pragma, for just that function?

Do you have an example? Is it like the __attribute specification?
Title: Re: GCC compiler optimisation
Post by: ataradov on August 30, 2022, 07:50:08 am
Yes, sure, if this only applies to one function. But given that -Og is used, it would not matter all that much. No need to micro-manage if you don't even macro-manage.

If -O3 causes "problems", you are very screwed and sitting on ticking time bomb. All such cases should be investigated immediately, not just ignored.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 30, 2022, 07:53:23 am
Do you have an example? Is it like the __attribute specification?
Just like any other attribute:

Code: [Select]
__attribute__ ((optimize ("no-tree-loop-distribute-patterns"))) void foo(void)
{
}
Title: Re: GCC compiler optimisation
Post by: peter-h on August 30, 2022, 08:21:24 am
Quote
If -O3 causes "problems", you are very screwed and sitting on ticking time bomb. All such cases should be investigated immediately, not just ignored.

IIRC this was stuff to do with memcpy etc substitution in the boot block which didn't have access to stdlib.

The reason I don't bother with it is because tests showed negligible code size or speed differences. As one would expect, the stuff got increasingly esoteric. Also some mentioned that -O3 ought to be considered "experimental" which is not really what I want to be doing.

The wider issue is that whenever something like this is changed, you have to do regression testing, which is impossible unless the product is trivial.

As I mention below, there are various places where one is relying on execution time to meet some hardware specs. How one does this, varies. If one needs to achieve say 1us (which is a looong time on a 168MHz 32F4) then a proper delay is necessary, and it needs to be scope-checked. It has to be a code loop (in asm, or in C with opt turned right off, or using CCYCNT). It can't be a "tick" because you would need a multi-MHz interrupt :) If one needs to achieve say 10ms, then osDelay(10) is usually the right way. These two are immune to global optimisation settings. But if you need say 50ns then what? You will probably get that with a dozen lines of code, and absolutely it must be scope-checked. And that one is vulnerable to -O settings being changed, unless the code is in a function with the -O0 attribute (which is what I have usually done).

So I don't agree that code which breaks with -O3 is necessarily bad code.

I am writing up a "design document" as I go along and these are my notes on this topic

Much time has been spent on this. It is a can of worms, because e.g. optimisation level -O3 replaces this loop
for (uint32_t i=0; i<length; i++)
{
 buf[offset+i]=data;
)
with memcpy() which while “functional” will crash the system if you have this loop in the boot loader and memcpy() is located at some address outside (which it will be, being a part of the standard library!) Selected boot block functions therefore have the attribute to force -O0 (zero optimisation) in case somebody tries -O3.

The basic optimisation of -O0 (zero) works fine, is easily fast enough for the job, and gives you informative single step debugging, but it produces about 30% more code. The next one, -Og, is the best one to use and doesn’t seem to do any risky stuff like above.

Arguably, one should develop with optimisation ON (say -Og) and then you will pick up problems as you go along. Then switch to -O0 for single stepping if needed to track something down. Whereas if you write the whole thing with -O0 and only change to something else later, you have an impossible amount of stuff to regression-test.

The problems tend to be really subtle, especially if timing issues arise. For example the 20ns min CS high time for the 45DB serial FLASH can be violated (by two function calls in succession) with a 168MHz CPU. A lot of ST HAL code works only because it is so bloated.

These figures show the relative benefits, at a particular stage of the project

-O0 produces 230k
-Og produces 160k
-O2 produces 160k*
-O3 produces 180k*
-Os produces 146k

The ones marked * are risky. -Os is pointless unless you are pushing against the FLASH size limit. Others, not listed above, have not been tested.

A compiler command line switch -fno-tree-loop-distribute-patterns has been added to prevent memcpy etc substitutions globally.


So yes probably -Og does not substitute stdlib functions.
Title: Re: GCC compiler optimisation
Post by: DiTBho on August 30, 2022, 08:33:49 am
compiler memory barriers and such constructs extremely rarely

that's one of the major reasons why I implemented myC: with tr-memory machines C doesn't offer *ANY* good memory barrier, and "Volatile" not only is futile but also a keyword that 90% of programmers don't understand and simply throw stuff at the wall, and see what sticks.
Title: Re: GCC compiler optimisation
Post by: DiTBho on August 30, 2022, 08:38:16 am
Do you have an example? Is it like the __attribute specification?
Just like any other attribute:

Code: [Select]
__attribute__ ((optimize ("no-tree-loop-distribute-patterns"))) void foo(void)
{
}

And then you have these solutions full of "black voodoo magic" (oh magic attribute, oh, what? see the manual, oh what? see the manual of YOUR C compiler, Oh SHT it's not supported, now what?) rather than neat stuff  :palm:

Title: Re: GCC compiler optimisation
Post by: DiTBho on August 30, 2022, 08:49:29 am
if you protect the access to shared objects [..] with mutexes, then you get implied memory barriers at the points where you acquire and release the mutex

Yes, XINU/R18200 (MIPS4+, multi core) has mutex an tr-memory primitives implemented in assembly in a dedicated module to avoid problems with the C Compiler, even because you need Pipeline special instructions for these operations.

Called "critical code". Small portion of assembly. Good compromise.
Title: Re: GCC compiler optimisation
Post by: gf on August 30, 2022, 09:57:42 am

that's one of the major reasons why I implemented myC: with tr-memory machines C doesn't offer *ANY* good memory barrier, and "Volatile" not only is futile but also a keyword that 90% of programmers don't understand and simply throw stuff at the wall, and see what sticks.

C basically does define an abstract, machine-independent memory ordering model, implemented via the stuff in stdatomic.h. Programs which strictly adhere to this abstract model (even if the current target CPU's requirements are not that strict) are even supposed to be portable to different CPUs, wrt. this functionality.
Title: Re: GCC compiler optimisation
Post by: DiTBho on August 30, 2022, 10:42:08 am
C basically does define an abstract, machine-independent memory ordering model, implemented via the stuff in stdatomic.h. Programs which strictly adhere to this abstract model (even if the current target CPU's requirements are not that strict) are even supposed to be portable to different CPUs, wrt. this functionality.

yeah, the concurrency support library, basically typedef _atomic (include/stdatomic.h)
Title: Re: GCC compiler optimisation
Post by: DiTBho on August 30, 2022, 11:26:53 am
support for this
Code: [Select]
enum memory_order
{
    memory_order_relaxed,
    memory_order_consume,
    memory_order_acquire,
    memory_order_release,
    memory_order_acq_rel,
    memory_order_seq_cst
};
is what is available since C11.

It's what I have always workarounded and segregated into assembly with projects to be compiled with previous C89 C compilers.
Title: Re: GCC compiler optimisation
Post by: gf on August 30, 2022, 12:10:10 pm

is what is available since C11.

It's what I have always workarounded and segregated into assembly with projects to be compiled with previous C89 C compilers.

Prior to that, gcc already suported the __sync_...() built-in functions. I think they were proprietary extensions, and most of them were full barriers.
Title: Re: GCC compiler optimisation
Post by: Nominal Animal on August 30, 2022, 12:26:37 pm

is what is available since C11.

It's what I have always workarounded and segregated into assembly with projects to be compiled with previous C89 C compilers.

Prior to that, gcc already suported the __sync_...() built-in functions. I think they were proprietary extensions, and most of them were full barriers.
They were originally introduced by Intel, in the Intel Itanium ABI, I believe.

Gcc has provided the __atomic_...() (https://gcc.gnu.org/onlinedocs/gcc-4.7.0/gcc/_005f_005fatomic-Builtins.html) built-in functions with the C++11 memory model since version 4.7 (2012), possibly earlier.
Title: Re: GCC compiler optimisation
Post by: ataradov on August 30, 2022, 03:05:59 pm
And then you have these solutions full of "black voodoo magic"
Optimizations are by definition compiler-specific. And there are like 2 useful modern compilers anyway.  And you don't have to do it though the attribute, you can just use a global flag. I hope you are not upset that different compilers have different command line flags?
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on August 30, 2022, 06:30:21 pm
If you have a problem with optimizations, it's 99.99% of the time because you are writing non-compliant code (consciously or not.)
My default optimization level with GCC and C is "-O3" for most embedded development I do, unless there is a problem with binary size, for which I'll usually switch to "-Os".
Never had a single issue related to this.

But I only use relatively recent targets, so nothing old, exotic or only supported by old or possibly unofficial compiler versions.
Title: Re: GCC compiler optimisation
Post by: DiTBho on August 30, 2022, 07:01:41 pm
Gcc has provided the __atomic_...() (https://gcc.gnu.org/onlinedocs/gcc-4.7.0/gcc/_005f_005fatomic-Builtins.html) built-in functions with the C++11 memory model since version 4.7 (2012), possibly earlier.

I think that even IBM has interests in supporting that stuff for their POWER10 and POWER11 and their tr-mem.

Gcc-v11 is very promising!

But I don't know, C++11&C look like ballet dancers; "Consume" is deprecated in C++17 because essentially nobody has been able to implement it in any way that's better than "acquire".

The first model of tr-mem was based on "Consume", now they say you can think of "Consume" as a restricted version of "Acquire", and "Relaxed" imposes no memory order at all.

You get the point  :-//
Title: Re: GCC compiler optimisation
Post by: DiTBho on August 30, 2022, 07:20:57 pm
(
ok, to tell you the true story, they said that C++17's design for "Consume" was "impractical" for Gcc to implement, so they gave up and strengthen it to "Acquire", but requiring a memory barrier on most weakly-ordered ISAs.

In real life, it's not a problem for x86 and ARM64, but it's a problem for PowerPC, MIPS4 (and more than horrible on MIPS4+) when to get that juicy performance devs need to actually use "Relaxed" but *very* carefully", more carefully than with x86 and ARM64, hoping it won't get optimized into something unsafe, because in this case it's really weak on weakly-ordered ISA therefore prone to catastrophic failures without the right enforced memory barriers.

So, you can see why I don't trust the C/C++ compiler and prefer the old school of segregating "critical code" into assembly modules.
)
Title: Re: GCC compiler optimisation
Post by: ejeffrey on August 30, 2022, 08:39:53 pm
The reason I don't bother with it is because tests showed negligible code size or speed differences. As one would expect, the stuff got increasingly esoteric. Also some mentioned that -O3 ought to be considered "experimental" which is not really what I want to be doing.

That isn't correct.  -O3 is not experimental.  All valid code should compile fine with -O3.  Of course compilers do have bugs but that is very rare.  Traditionally -O3 enables optimizations which have a reasonable chance to make performance worse: frequently ones that can cause a large increase in compiled code size (which can reduce pipeline stalls but increase cache pressure).

If your code breaks with -O3 then it is almost always non-compliant.

Some compilers have other options that relax constraints and can generate non-conformant behavior but perform even faster if you know this is OK for your application.  But those options should never be turned on by -O?.
Title: Re: GCC compiler optimisation
Post by: gf on August 30, 2022, 09:35:50 pm
Quote
ok, to tell you the true story, they said that C++17's design for "Consume" was "impractical" for Gcc to implement, so they gave up and strengthen it to "Acquire", but requiring a memory barrier on most weakly-ordered ISAs.

In real life, it's not a problem for x86 and ARM64, but it's a problem for PowerPC, MIPS4 (and more than horrible on MIPS4+) when to get that juicy performance devs need to actually use "Relaxed" but *very* carefully", more carefully than with x86 and ARM64, hoping it won't get optimized into something unsafe, because in this case it's really weak on weakly-ordered ISA therefore prone to catastrophic failures without the right enforced memory barriers.

It is of course always safe to implement Consume as Acquire. My feeling is that Consume was "invented" in order to have a faster alternative for some special cases of Acquire, for those processors which require an (inperformant) memory barrier instruction for the implementation of Acquire semantics.

On x86, regular loads have Acquire semantics per se, so no extra instructions are necessary in order to obtain Consume or Acquire semantics, and load-Consume and load-Acquire degrade to compiler-only barriers which only prevent instruction re-ordering by the compiler.

EDIT:
Btw, here's an interesing article about Consume:
https://preshing.com/20140709/the-purpose-of-memory_order_consume-in-cpp11/
Title: Re: GCC compiler optimisation
Post by: peter-h on February 09, 2023, 05:08:50 pm
I've been digging around memcpy lately (because it is used in the ST 32F4 ETH PHY low_level_* code) and having found the default Newlib memcpy does just one byte at a time, I switched to this one, which is also from Newlib but optimised for speed

Code: [Select]

#define BIGBLOCKSIZE    (sizeof (long) << 2)
#define LITTLEBLOCKSIZE (sizeof (long))
#define TOO_SMALL(LEN)  ((LEN) < BIGBLOCKSIZE) /* Threshhold for punting to the byte copier.  */
/* Nonzero if either X or Y is not aligned on a "long" boundary.  */
#define UNALIGNED(X, Y) (((long)X & (sizeof (long) - 1)) | ((long)Y & (sizeof (long) - 1)))

void * memcpy (void *__restrict dst0, const void *__restrict src0, size_t len0)
{

char *dst = dst0;
const char *src = src0;
long *aligned_dst;
const long *aligned_src;

/* If the size is small, or either SRC or DST is unaligned,
     then punt into the byte copy loop.  This should be rare.  */
if (!TOO_SMALL(len0) && !UNALIGNED (src, dst))
{
aligned_dst = (long*)dst;
aligned_src = (long*)src;

/* Copy 4X long words at a time if possible.  */
while (len0 >= BIGBLOCKSIZE)
{
*aligned_dst++ = *aligned_src++;
*aligned_dst++ = *aligned_src++;
*aligned_dst++ = *aligned_src++;
*aligned_dst++ = *aligned_src++;
len0 -= BIGBLOCKSIZE;
}

/* Copy one long word at a time if possible.  */
while (len0 >= LITTLEBLOCKSIZE)
{
*aligned_dst++ = *aligned_src++;
len0 -= LITTLEBLOCKSIZE;
}

/* Pick up any residual with a byte copier.  */
dst = (char*)aligned_dst;
src = (char*)aligned_src;
}

while (len0--)
*dst++ = *src++;

return dst0;

}


I am seeing a speedup of about 20% (time across the whole function, on packets of a few hundred bytes) which is a lot less than I would expect. Is that code really meaningful for an arm32? I can see what it is doing, I think. The typical exec time is 15us which is probably not worth optimising further (e.g. with DMA).

Could anyone also advise why the "const" in
const long *aligned_src;

Thank you.
Title: Re: GCC compiler optimisation
Post by: NorthGuy on February 09, 2023, 05:24:03 pm
optimised for speed

Not by much. Agner Fog wrote a very good book on x86 optimizatons, which includes lots of interesting ideas on memory/string routine optimizations. x86 is an out-of-order CPU, so it's not completely transferable to small ARM, but you can read the book, see what's involved, and write optimized routines if you have a desire and time for this.
Title: Re: GCC compiler optimisation
Post by: Siwastaja on February 09, 2023, 05:29:36 pm
Could anyone also advise why the "const" in
const long *aligned_src;

As the name suggests, const in C language communicates (both to the compiler, and human reader) that the thing is constant - it does not change.

Pretty obviously, given the purpose of memcpy, the source memory buffer is not changed, only destination is. The idea is to use the const keyword, so that if compiler detects someone is assigning to a const qualified object, it will be illegal so the compiler will raise an error and catch a dangerous bug.

Improve your code quality and make it a habit to always qualify things that are not supposed to change const.

Note that C declarations are read from right to left. Here, the actual memory is qualified const. Pointer is not const, so the pointer can be changed (to point somewhere else, like the next word). If one wants to declare a const pointer to const memory, that would be:
const long * const aligned_src; or equivalently:
long const * const aligned_src;
Title: Re: GCC compiler optimisation
Post by: peter-h on February 09, 2023, 05:31:54 pm
The biggest optimisation should be moving 32 bits at a time, which should be simply 4x faster, on a decent size block. The 32F4 has no data cache.

And I am using a -O3 attribute.

Well, unless this function is trying to do 32 bit unaligned moves, which is dumb, even though the 32F4 does support unaligned access.

So maybe my src or dest are unaligned, in which case the 32 bit loop is skipped, but what would be dumb code. The smart way is to move unaligned bytes first, then move 32 bits at a time, and then clean up with byte moves.

I ought to look at this
https://cboard.cprogramming.com/c-programming/154333-fast-memcpy-alternative-32-bit-embedded-processor-posted-just-fyi-fwiw.html
but would have expected a lot less code.

Title: Re: GCC compiler optimisation
Post by: brucehoult on February 09, 2023, 08:12:13 pm
x86 is an out-of-order CPU, so it's not completely transferable to small ARM

x86 is an instruction set. Some implementations of x86 are out of order (ok, most these days), and some are not. This one, for example, is in-order:

https://www.digikey.co.nz/en/products/detail/rochester-electronics-llc/N80186/12122323 (https://www.digikey.co.nz/en/products/detail/rochester-electronics-llc/N80186/12122323)

ARM is an instruction set (3 or 4 different ones, actually). Some implementations of ARM are out of order (A15 and A57+), and some are not.

OoO gives you a bit more latitude in scheduling your instructions, and especially in not having to unroll loops, but the biggest issue in optimisation is issue/execute width. There are 3-wide in-order CPUs (including from ARM). There are 2-wide microcontrollers.
Title: Re: GCC compiler optimisation
Post by: peter-h on February 09, 2023, 08:22:59 pm
I've just realised that this isn't possible to do fully, because if you start with the two buffers being unaligned, but unaligned differently, then no number of single byte transfers is going to get you to a point where you could continue with 32 bit transfers. So if e.g. one starts at 0x10000001 and the other at 0x10000002, you have to do the whole copy in the byte mode.

One option might be to use the 32 bit mode anyway because the 32F4 supports unaligned 32 bit transfers, until there are just 0-3 bytes left and then continue in a single byte mode. Then you will automatically get aligned transfers if the two buffers are both aligned, automatically.

I also found that the four consecutive 32 bit moves were always getting optimised out, and only one was executed, and the compiler was decrementing the loop counter appropriately :)

So I now have

Code: [Select]

#define BLOCKSIZE    (sizeof (long))
#define TOO_SMALL(LEN)  ((LEN) < BLOCKSIZE) /* Threshhold for punting to the byte copier.  */
/* Nonzero if either X or Y is not aligned on a "long" boundary.  */
#define UNALIGNED(X, Y) (((long)X & (sizeof (long) - 1)) | ((long)Y & (sizeof (long) - 1)))

__attribute__((optimize("O2")))
void * memcpy_fast (void *__restrict dst0, const void *__restrict src0, size_t len0)
{

char *dst = dst0;
const char *src = src0;
long *aligned_dst;
const long *aligned_src;

/* If the size is small, or either SRC or DST is unaligned,
     then punt into the byte copy loop.  This should be rare.  */
if (!TOO_SMALL(len0) && !UNALIGNED (src, dst))
{
aligned_dst = (long*)dst;
aligned_src = (long*)src;

/* Copy long words if possible.  */
while (len0 >= BLOCKSIZE)
{
*aligned_dst++ = *aligned_src++;
len0 -= BLOCKSIZE;
}

/* Pick up any residual with a byte copier.  */
dst = (char*)aligned_dst;
src = (char*)aligned_src;
}

while (len0--)
*dst++ = *src++;

return dst0;

}

Can't get my head around how to do the pointer notation for my earlier proposal (always 32 bit until < 4 bytes left).
Title: Re: GCC compiler optimisation
Post by: Nominal Animal on February 09, 2023, 08:39:15 pm
The proper test is actually (((uintptr_t)src) ^ ((uintptr_t)dst) & 3).  It is zero whenever 32-bit aligned access is possible, and nonzero when the copy is inherently unaligned (ie. even when src is aligned, dst is unaligned, and vice versa).

When the above expression is zero, (((uintptr_t)src) & 3) tells you the number of leading bytes that need to be transferred before an aligned transfer is possible.  If your MCU supports unaligned accesses, then doing a single unaligned 32-bit copy, followed by aligning src and dst to the next 32-bit boundary (so between one and three bytes will be transferred again), is likely to be more efficient than eg. Duff's Device bytewise copy of the initial bytes.

Similarly, if there are any trailing bytes, doing a single unaligned 32-bit transfer to copy the trailing bytes, is probably more efficient than Duff's Device bytewise copy or a trailing copy loop, whenever unaligned accesses are supported (because the extra cost tends to be single digit cycles).

Of course, if the unaligned access cost is low enough, then just explicitly aligning one (the destination, assuming that it will be more likely to be accessed next rather than the source, caching-wise) and letting the other be unaligned, is even better because the test/selection logic before the copy is done has a significant cost compared to typical memcpy() sizes.

For explicit calls where you know the pointers and the data are 32-bit aligned –– even via
    if (_Alignof(*src) >= 4 && _Alignof(*dst) && !(n & 3))
        memcpy32p((uint32_t *)dst, (const uint32_t *)src, (uint32_t *)dst + ((uint32_t)n / 4));
    else
        memcpy8p((unsigned char *)dst, (const unsigned char *)src, (unsigned char *)dst + n);
where the third parameter is a the limit to dst, and not the number of units to copy –– for example as an inline wrapper around your copy function, can shave off the extra test/selection logic cost.
Title: Re: GCC compiler optimisation
Post by: eutectique on February 09, 2023, 08:44:30 pm
You can ensure 4-byte alignment of memory regions with __attribute__ ((aligned (4))).
Title: Re: GCC compiler optimisation
Post by: peter-h on February 09, 2023, 09:07:21 pm
Interesting.

Indeed I take care to align buffers, but in some cases (ones inside LWIP) I can't possibly know what the addresses will be if one is picking up some buffer part way through.

Quote
then just explicitly aligning one

I didn't think of that... one being aligned is better than neither.

But not checking for alignment seems to automatically yield the optimal solution - on the 32F4:

Code: [Select]

// Optimised memcpy. This is based on one in Newlib but specifically for the 32F4
// which does unaligned 32 bit transfers transparently. This avoids having to check
// buffer alignment, and 4-aligned buffers are automatically optimised internally.

__attribute__((optimize("O2")))
void * memcpy_fast (void *__restrict dst0, const void *__restrict src0, size_t len0)
{

char *dst = dst0;
const char *src = src0;
uint32_t *aligned_dst;
const uint32_t *aligned_src;

// If the size is >=4 then do 32 bit moves, until exhausted

if ( len0 >=4 )
{
aligned_dst = (uint32_t*)dst;
aligned_src = (uint32_t*)src;

while (len0 >= 4)
{
*aligned_dst++ = *aligned_src++;
len0 -= 4;
}

dst = (char*)aligned_dst;
src = (char*)aligned_src;
}

// Finish with any single byte moves

while (len0--)
*dst++ = *src++;

return dst0;

}

I am just not 100% sure the above code cannot run off the end.

I am now getting a 2x speedup over the byte version, on average.

I started on a DMA version but quickly remember that it would be no good due to lack of CCM access. A lot of the time I can explicitly control this (like using DMA for all SPI transfers) but with LWIP, who knows. I have the RTOS stacks in CCM.

Title: Re: GCC compiler optimisation
Post by: ataradov on February 09, 2023, 09:41:51 pm
I started on a DMA version but quickly remember that it would be no good due to lack of CCM access.
You can check pointer ranges and figure out if both pointers are in a DMA region, then use DMA, otherwise use slower version.
Title: Re: GCC compiler optimisation
Post by: brucehoult on February 09, 2023, 09:51:09 pm
The proper test is actually (((uintptr_t)src) ^ ((uintptr_t)dst) & 3).  It is zero whenever 32-bit aligned access is possible, and nonzero when the copy is inherently unaligned (ie. even when src is aligned, dst is unaligned, and vice versa).

When the above expression is zero, (((uintptr_t)src) & 3) tells you the number of leading bytes that need to be transferred before an aligned transfer is possible.  If your MCU supports unaligned accesses, then doing a single unaligned 32-bit copy, followed by aligning src and dst to the next 32-bit boundary (so between one and three bytes will be transferred again), is likely to be more efficient than eg. Duff's Device bytewise copy of the initial bytes.

Even if your CPU supports unaligned accesses, using that capability might be slower than using a loop like the following (plus setup and cleanup of a few bytes at start and end). byteOffset is 0,1,2,3. Assuming little-endian.

Code: [Select]
uint32_t copyLoop(uint32_t *src, uint32_t *dst, int byteOffset, uint32_t *srcLimit, uint32_t dstBytes){
    int bitOffset = byteOffset<<3;
    while (src < srcLimit){
        uint32_t srcBytes = *src++;
        dstBytes |= srcBytes<<bitOffset;
        *dst++ = dstBytes;
        dstBytes = srcBytes >> (32-bitOffset);
    }
    return dstBytes;
}

On ARM the body of the loop is:

Code: [Select]
        rsb     rightShift, bitOffset, #32
loop:
        ldr     srcBytes, [src], #4
        orr     dstBytes, dstBytes, srcBytes, lsl bitOffset
        str     dstBytes, [dst], #4
        lsr     dstBytes, srcBytes, rightShift
        cmp     srcLimit, src
        bhi     loop

Of course you can unroll / Duff this as you wish, but it probably goes at close to full cache speed anyway.

NB: taking this literally is UB in C for the right shift with byteOffset=0. It works fine on ARM (with register source for the shift amount) but not on others such as x86 and RISC-V where a shift of 32 is a shift of 0. You can make a separate simple *dst++ = *src++ loop for the byteOffset==0 case.
Title: Re: GCC compiler optimisation
Post by: peter-h on February 09, 2023, 10:13:46 pm
Yes; I did that in some other code. Use DMA if not in CCM. 2 questions:

1) Would DMA beat aligned 32 bit memcpy? It might do because one can do the same "auto unaligned support" hack on DMA. One would still need to finish off with software copy unless blocksize is a multiple of 4.

2) Any gotchas for DMA stream for memory-memory? I remember reading that most of the DMA streams cannot be used.

It's also a good point that even in the case of both buffers being unaligned, and even if they are unaligned differently, it will be faster to copy single bytes until one of the buffers becomes aligned, and then 32 bit moves will run faster.
Title: Re: GCC compiler optimisation
Post by: ataradov on February 09, 2023, 11:19:19 pm
DMA will only do aligned transfers. But DMA may be fast enough to just transfer bytes and not worry about further optimization. I think there are also smart features that would pack multiple bytes into a word, but I don't know the details, and don't know if this is something useful here.

No idea about limitations of the specific device, but it should not be too hard to figure out.
Title: Re: GCC compiler optimisation
Post by: peter-h on February 10, 2023, 07:15:48 am
If DMA does aligned only (I didn't know that) it will struggle to beat a software copy which will benefit from alignment a lot of the time.
Title: Re: GCC compiler optimisation
Post by: ataradov on February 10, 2023, 07:18:07 am
It has independent access to two interfaces. There is no scenario where the most optimized aligned version doing word transfers in the firmware beats simple byte transfers using DMA.
Title: Re: GCC compiler optimisation
Post by: peter-h on February 10, 2023, 10:10:04 am
Are memory-memory DMA transfers interruptible?

I know memory-peripheral ones obviously are - in between bytes moved. But there the DRQ stream has large gaps. In memory-memory you would try to have no gaps.
Title: Re: GCC compiler optimisation
Post by: Siwastaja on February 10, 2023, 10:55:40 am
Are memory-memory DMA transfers interruptible?

I know memory-peripheral ones obviously are - in between bytes moved. But there the DRQ stream has large gaps. In memory-memory you would try to have no gaps.

Why not RTFM? For example from an STM32H7 reference manual:

"At any time, a DMA transfer can be suspended to be restarted later on or to be definitively
disabled before the end of the DMA transfer.
There are two cases:
Note:
• The stream disables the transfer with no later-on restart from the point where it was
stopped. There is no particular action to do, except to clear the EN bit in the
DMA_SxCR register to disable the stream. The stream may take time to be disabled
(ongoing transfer is completed first)."

It says nothing about memory-to-memory transfers being special, so very likely it respects the EN bit.

But the question is, why would you want to do that?
Title: Re: GCC compiler optimisation
Post by: gf on February 10, 2023, 11:19:56 am
The question is rather how much memory bandwidth is consumed by the DMA transfer, and is there enough left for the concurrent duties of the CPU.
Title: Re: GCC compiler optimisation
Post by: peter-h on February 10, 2023, 11:43:02 am
Quote
why would you want to do that?

If a DMA transfer takes say 100us, that is the interrupt service latency.

Not too bad if a byte is really moved on every clock. A max-size ETH packet works out to 10us. But I am already seeing memcpy() times in that approximate range. Never more than 20us. Can't really correlate them to packet sizes (would have to waggle some other GPIO based on the "length" value) but I don't see DMA as being much faster. For real speedups on LWIP one needs to move to zero-copy and that seems to be a can of worms (even though it is a pretty old idea with LWIP).

Quote
The question is rather how much memory bandwidth is consumed by the DMA transfer, and is there enough left for the concurrent duties of the CPU.

Indeed.

The 32F4 manual doesn't say anything about mem-mem transfers being interruptible. Perhaps burst transfers might be but they must not cross a 1k address boundary, which is useless.

This is interesting because the 1970s Z80-family chips could do that :) It was called "cycle stealing". Presumably you could achieve the same by using a timer to trigger each transfer, but then the entire transfer would obviously be slower. But this would be a way to achieve a slow background transfer process.
Title: Re: GCC compiler optimisation
Post by: peter-h on February 10, 2023, 03:42:57 pm
On the topic of bizzare optimisation...

This

Code: [Select]
__attribute__((optimize("O2")))
void * memcpy_fast (void *__restrict dst0, const void *__restrict src0, size_t len0)
{

char *dst = dst0;
const char *src = src0;
uint32_t *aligned_dst;
const uint32_t *aligned_src;

// If the size is >=4 then do 32 bit moves, until exhausted

if ( len0 >=4 )
{
aligned_dst = (uint32_t*)dst;
aligned_src = (uint32_t*)src;

while (len0 >= 4)
{
*aligned_dst++ = *aligned_src++;
len0 -= 4;
}

dst = (char*)aligned_dst;
src = (char*)aligned_src;
}

// Finish with any single byte moves

while (len0--)
*dst++ = *src++;

return dst0;

}

when stepped through, reveals

(https://peter-ftp.co.uk/screenshots/20230210475223915.jpg)

One has to laugh, given that "prevailing internet wisdom" has been that only -O3 and above does auto loop identification, and I have

(https://peter-ftp.co.uk/screenshots/20230210525234115.jpg)

I think some compiler writers need to get a life. They must have spent years doing these dumb stdlib substitutions.
Title: Re: GCC compiler optimisation
Post by: ataradov on February 10, 2023, 04:58:25 pm
If a DMA transfer takes say 100us, that is the interrupt service latency.
Of course they are interruptible. DMA controller negotiates access for each transfer just like any other master. Individual transfers within the burst transfer are atomic.

Quote
The question is rather how much memory bandwidth is consumed by the DMA transfer, and is there enough left for the concurrent duties of the CPU.
Indeed.
You are doing blocking memcpy(). How much concurrent stuff do you expect to happen?

Perhaps burst transfers might be but they must not cross a 1k address boundary, which is useless.
This is a strange logic. Burst transfers are not interrtuptible. The whole point of the burst transfer is to negotiable access once and perform it atomically.

And why is 1K useless? Just split your buffer into 1K chunks and do individual transfer for each. But this is not necessary.

As usual, you are way overthinking all this.
Title: Re: GCC compiler optimisation
Post by: westfw on February 10, 2023, 05:14:23 pm
DMA transfers aren’t usually said to be “interruptible”, but rather “stallable”. Each memory transfer (or perhaps each “burst”) is subject to having to wait before it is granted access to memory.
On a risc platform where the cpu can probably access ram nearly as fast as it can go, this might be more of a problem than elsewhere.  Frequently there are priorities involved (sometimes user settable?)

Title: Re: GCC compiler optimisation
Post by: peter-h on February 10, 2023, 05:53:41 pm
Quote
You are doing blocking memcpy()

Blocking yes but ints are serviced as normal, RTOS can switch tasks, etc.

Can anyone explain why the call to the library memcpy() was there? We had a long thread on this ages ago. I took a lot of care to avoid this happening in my boot block (which has no access to stdlib) and it was done with a combination of -O0 function attribute and that distribute-patterns compiler option, but the latter ought to have been sufficient.
Title: Re: GCC compiler optimisation
Post by: brucehoult on February 10, 2023, 06:14:26 pm
Can anyone explain why the call to the library memcpy() was there?

Because the compiler is allowed to replace code such as...

Code: [Select]
while (len0--)
  *dst++ = *src++;

... with memcpy.

And, yes, '-fno-tree-loop-distribute-patterns' should stop it doing that.
Title: Re: GCC compiler optimisation
Post by: ataradov on February 10, 2023, 06:55:58 pm
On a risc platform where the cpu can probably access ram nearly as fast as it can go, this might be more of a problem than elsewhere.  Frequently there are priorities involved (sometimes user settable?)
In case of STM32F4xx bus matrix uses fixed round-robin arbitration, so all masters that want to get access would be granted it with a fixed priority in turn.  But yes, in many cases priorities and scheduling algorithm are configurable.

And the matrix itself is multi-layer, so multiple masters may get access at the same time as long as they are accessing different slaves. It may not help as much in case of SRAM access, but even here SRAM is split into 3 blocks, which may be accessed independently. So, if you are really deep into optimizing stuff, you can shuffle the buffers around the memory. In practice all this hits diminishing returns real fast. If you have to do this - get a faster CPU.
Title: Re: GCC compiler optimisation
Post by: Siwastaja on February 10, 2023, 07:26:46 pm
If a DMA transfer takes say 100us, that is the interrupt service latency.

Wait, wat?!

Where do you get these ideas? I mean, I have never ever seen someone come up with stuff like that. That would render DMA completely useless.

No. That's not how it works. Bus arbiter gives the CPU memory access cycles. If my memory serves right, for STM32, it's at least every other memory cycle (see the documentation for exacts).
Title: Re: GCC compiler optimisation
Post by: peter-h on February 10, 2023, 08:44:15 pm
Quote
Where do you get these ideas? I mean, I have never ever seen someone come up with stuff like that

As you know, I am stupid, so treat this as yet another opportunity to show me that you are much more clever ;) Seriously, I really like working with people much smarter than I am, because every day is a school day (an English saying, but I am not actually English).

I could not find an answer to whether mem-mem DMA transfers are interruptible in the RM.

Quote
And, yes, '-fno-tree-loop-distribute-patterns' should stop it doing that.

Yes, it should, but for some reason it doesn't.
Title: Re: GCC compiler optimisation
Post by: ataradov on February 10, 2023, 08:47:35 pm
Yes, it should, but for some reason it doesn't.
-fno-builtin would disable use of any compiler-builtin version of the libc functions. This would disable this optimization, since compiler can no longer rely in memcpy() being there.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on February 10, 2023, 08:56:04 pm
Yes, it should, but for some reason it doesn't.
-fno-builtin would disable use of any compiler-builtin version of the libc functions. This would disable this optimization, since compiler can no longer rely in memcpy() being there.

Not sure why you'd want to do that though.
Title: Re: GCC compiler optimisation
Post by: ataradov on February 10, 2023, 09:09:01 pm
Not sure why you'd want to do that though.
May be useful in freestanding code, but there is a separate  flag for this anyway.

And on a closer look it is also not that. Even though it seems to work as a side effect, -fno-builtin is meant to disable replacement of the library functions with internal code, so the opposite of what happens here.

In this case some other optimization that recognizes memcpy semantics kicks in. The call is external anyway in this case.

-fno-builtin-memcpy, which should only disable memcpy, does not do anything here. More investigation is required.

If you want to experiment - https://godbolt.org/z/cx1GPxEE1

And -fno-tree-loop-distribute-patterns does work as far as I can see. But I don't think is guaranteed, since this does not actually disable replacement with library calls, it just disables optimizations that can allow replacements.

And upon further reading, -ftree-loop-distribute-patterns is the optimization that does this. So, it should always work.
Title: Re: GCC compiler optimisation
Post by: peter-h on February 10, 2023, 09:26:24 pm
Quote
May be useful in freestanding code, but there is a separate  flag for this anyway.

What flag is that?

That's exactly what I am doing (normally, e.g. a boot loader).

But also I wrote a faster replacement for memcpy, and the compiler recognised the loop and replaced it with a slower memcpy! That's just daft, and it got it wrong!

A -O0 function attribute does seem to work ok in suppressing these substitutions, at the cost of ~ 1/3 more code.
Title: Re: GCC compiler optimisation
Post by: ataradov on February 10, 2023, 09:30:02 pm
What flag is that?
Literally -ffreestanding and you may also want -fnostdlib

But -fno-tree-loop-distribute-patterns should disable that replacement too. If it does not, I'd investigate that. Something is not right in your setup. Maybe the flag does not actually get to the compiler or something.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on February 10, 2023, 09:38:46 pm
Not sure why you'd want to do that though.
May be useful in freestanding code, but there is a separate  flag for this anyway.

There are two separate things here. While you may want to have no dependency on the std lib at all, which would preclude any call to a memcpy/memset/... , if the compiler's optimizer actually inlines some memcpy-like code without any function call, then technically is it using the std lib or not? And why would you want to prevent the compiler from doing that, while the resulting code would be pretty much the same as directly compiling your hand-written copy, except usually better.
Title: Re: GCC compiler optimisation
Post by: peter-h on February 10, 2023, 10:06:10 pm
See the example code I posted.

It takes advantage of 25% of moves being fully aligned (src+dest), 50% of moves being partly aligned (src or dest), and (in the 32F4) still works with 25% of moves being fully unaligned.

Whereas the code in Newlib is suboptimal 75% of the time.

In the specific case of memcpy this is more complicated because while the byte-only memcpy is from Newlib, the full memcpy in one Newlib source I found
https://github.com/eblot/newlib/blob/master/newlib/libc/string/memcpy.c
does have the fuller source which does the long optimisation. My faster memcpy was based on that.

My current settings are below and thus my byte-only memcpy suggests that ST compiled that library for minimum size, but I think, from other stuff I found, only Newlib-nano should be doing that, but my settings here are obviously not going to change a precompiled library.

(https://peter-ftp.co.uk/screenshots/202302105716515521.jpg)

The flag is being used on the command line

(https://peter-ftp.co.uk/screenshots/202302103516520522.jpg)

Quote
And why would you want to prevent the compiler from doing that

In this case, the compiler isn't smart enough to know what runs faster on that particular CPU. I can understand why it called a sub-optimal memcpy but can I assume that it would replace am optimised memcpy with an even more optimised one? I don't think so.
Title: Re: GCC compiler optimisation
Post by: ataradov on February 10, 2023, 10:45:53 pm
Library memcpy() is whatever it is. But if you just write your version in plain c and let compiler substitute internal implementation, you will get the same exact optimized version.

Here is test code:
Code: [Select]
void *memcpy_my(void *restrict dest, const void *restrict src, int len)
{
    char *dp = (char *restrict)dest;
    const char *sp = (const char *restrict)src;
    while( len-- )
    {
        *dp++ = *sp++;
    }
    return dest;
}

If you compiler this with -Og, then it would result in a simple byte by byte loop. But if you use real optimization levels, then it generates the same code that checks alignment and does whatever is necessary to take care of the edge conditions:

arm-none-eabi-gcc -O3 -mthumb -mcpu=cortex-m4 -fno-tree-loop-distribute-patterns -mno-unaligned-access  -c test.c -o test

And if you remove -mno-unaligned-access, then it will generate a simpler version of the code that uses unaligned word access, which is still plenty fast and not worth thinking about in practice.
Title: Re: GCC compiler optimisation
Post by: peter-h on February 10, 2023, 11:06:35 pm
Quote
But if you just write your version in plain c and let compiler substitute internal implementation, you will get the same exact optimized version.

Evidently not.

Quote
If you compiler this with -Og, then it would result in a simple byte by byte loop

Yes.

Quote
But if you use real optimization levels, then it generates the same code that checks alignment and does whatever is necessary to take care of the edge conditions

Evidently not if it calls a library memcpy. But maybe Yes if the library happened to have been built with an optimised version.

I did a google on -mno-unaligned-access and it is a big can of worms.

Quote
not worth thinking about in practice.

That would be generally true, but only because most people write some code and if it works they move on, so they don't discover what is going on ;)

Title: Re: GCC compiler optimisation
Post by: ataradov on February 10, 2023, 11:13:30 pm
Evidently not if it calls a library memcpy. But maybe Yes if the library happened to have been built with an optimised version.
It only calls the library in your code where you already attempt to optimize everything.

If you take my code from above and use it as is with the flags I provided, it should generate optimized code that does not call the library. If it does not - something is broken in your compiler.

What compiler version you are using anyway?

I did a google on -mno-unaligned-access and it is a big can of worms.
-munaligned-access may be dangerous, but -mno-unaligned-access is always safe.

If you want it to be absolutely optimal, use DMA.
Title: Re: GCC compiler optimisation
Post by: peter-h on February 10, 2023, 11:18:52 pm
Quote
What compiler version you are using anyway?

The latest with latest Cube. V10.something.

Not sure where it is easily visible, short of running one of the executables on a command line.
Title: Re: GCC compiler optimisation
Post by: ataradov on February 10, 2023, 11:25:12 pm
All the experiments above were done with "gcc version 10.2.1 20201103 (release) (GNU Arm Embedded Toolchain 10-2020-q4-major)" (official ARM release).

Works here as expected.
Title: Re: GCC compiler optimisation
Post by: Siwastaja on February 11, 2023, 06:55:16 am
As you know, I am stupid,

I don't think so at all. This is a perfect example. Normal stupid people don't think about DMA vs. CPU bus arbitration at all, and everything just works out because the engineers at ARM / ST did a sane design, for those stupid people to use.

Instead, you are smart enough to come up with a unique scenario. You are not stupid, you are overthinking and if anything, fixating into things that should be ruled out quickly, for too long, which eats time you could use for more fruitful thinking.

I suggest you learn to instrument your code. That way, you can always do a simple test/measurement when you feel like reading the documentation is not efficient use of time.

Quote
I could not find an answer to whether mem-mem DMA transfers are interruptible in the RM.

This is because you came up with your own search term which does not make sense ("interruptible"). Kinda X-Y problem. Your real question is, does using DMA slow down CPU doing normal CPU things (e.g., serve interrupts), and if it does, by how much. But if you just say "is DMA interruptible", no one understands what you are asking. I mean, the whole point of DMA is that it runs parallel with the CPU. That's literally the only raison d'etre. So usually there are very few cases in which you want to pause or cancel the DMA operation.

For example, this is explained on the Wikipedia page about DMA:
Quote
Direct memory access (DMA) is a feature of computer systems that allows certain hardware subsystems to access main system memory independently of the central processing unit (CPU).
and
Quote
With DMA, the CPU first initiates the transfer, then it does other operations while the transfer is in progress
Title: Re: GCC compiler optimisation
Post by: brucehoult on February 11, 2023, 07:12:12 am
I don't think so at all. This is a perfect example. Normal stupid people don't think about DMA vs. CPU bus arbitration at all, and everything just works out because the engineers at ARM / ST did a sane design, for those stupid people to use.

The same situation applies, after all, to multiple CPUs sharing a memory system.

A DMA unit is just a special purpose CPU with a very limited instruction set.
Title: Re: GCC compiler optimisation
Post by: Nominal Animal on February 11, 2023, 07:35:17 am
Words are hard, too.  Well, actually fuzzy and vague and not well defined or "solid" at all.

You also haven't met stupid until you meet a student who excitedly and happily tells you they managed to chop off all their fingers, so they wouldn't have to go to shop class, and could now concentrate better on their dream: becoming a professional concert pianist.
Title: Re: GCC compiler optimisation
Post by: peter-h on February 11, 2023, 07:46:11 am
Quote
because you came up with your own search term which does not make sense ("interruptible")

Hmm, no.

Quote
I suggest you learn to instrument your code

Hmm, have been doing that for 40+ years.
Title: Re: GCC compiler optimisation
Post by: wek on February 11, 2023, 08:23:24 am
The STM32 dual-port DMA used in STM32F4xx is described in AN4031. Chapter 2 deals with both arbitration at the bus matrix, and arbitration within the DMA itself.

Both are round-robin, so DMA won't hog the system.

There's also timing information there. M2M is nothing special, it works exactly the same as P2M, except the triggers are provided internally so that when one transfer finishes the next is started (in fact you can achieve the same effect by triggering transfers from sources which are not cleared by the transfer, e.g. using a SPI-RXNE-triggered transfer which does read from given SPI's data register (how do I know that...)). In ideal situation, one transfer would IMO take 3 or 4 cycles (I'm not ARM/ST insider, not exactly sure about the arbitration delays) if source and destination lay in different memories. The ideal situation should be easy to benchmark. How far real world is from the ideal situation, is of course dependent on dozens of pesky details.

JW
Title: Re: GCC compiler optimisation
Post by: peter-h on February 11, 2023, 09:11:22 am
Quote
one transfer would IMO take 3 or 4 cycles

That's interesting, since it suggests that software is no slower, especially if you do 4 bytes in one go.

DMA, byte at a time, at 168MHz, I would expect to be similar to an optimised loop.

Getting back to optimisation, I found that even -Og (my default mode for the whole project, with -O0 used for some functions) removes the four consecutive moves found in some memcpy sources. Stepping through assembler, I had one case where 2 of the 4 were removed (and the loop counter halved). Quite weird. I think to get a real unrolled loop you would need to use -O0 and perhaps "register" on some values.
Title: Re: GCC compiler optimisation
Post by: ataradov on February 11, 2023, 09:23:45 am
How are you doing copy in software 3 or even 4 cycles? Load and store instructions on CM4 take 2 cycles. Then you need to decrement the counter and branch, which is 2-4 cycles if taken, since it does pipeline flush and there is no branch predictor in CM4. And this is for zero wait state code memory or flash accelerator working out of its mind.

Note that consecutive load/store instructions of the same type may pipeline, so if you are going to do optimized software version, it is better to do 4 or 8 consecutive loads and then 4 or 4 stores. This amortizes the loop counter decrement and branch penalty.

DMA will always be faster. Also if you transfer between different SRAMx slaves, then it would pipeline and provided no other masters want to access the same SRAM, it would be 1 transfer/cycle sustained. This is not a likely scenario though, especially placement of buffers in different SRAMs.

I don't personally want to participate in optimization stuff anymore. I can't take that anymore.
Title: Re: GCC compiler optimisation
Post by: Nominal Animal on February 11, 2023, 10:10:53 am
How are you doing copy in software 3 or even 4 cycles? Load and store instructions on CM4 take 2 cycles.
Consider
Code: [Select]
loop:
    ldr r0, [r1], #4 ; 1 cycle (pipelined with following str)
    str r0, [r2], #4 ; 2 cycles
    cmp r1, r3       ; 1 cycle
    blt loop         ; 2 cycles
Each loop iteration takes 6 cycles, right?  As described in Cortex M4 Technical Reference Manual, 3.3.2 Load/Store Timings.  (I might be wrong, though, because it does not explicitly describe the address post-increment case, [register],#increment, and only refers to the [register,offset] forms, with paired ldr-str taking three cycles.)

If the loop is unwound, say
Code: [Select]
loop:
    ldr r0, [r1], #4 ; 1 cycle (pipelined with following str)
    str r0, [r2], #4 ; 2 cycles
    ldr r0, [r1], #4 ; 1 cycle (pipelined with following str)
    str r0, [r2], #4 ; 2 cycles
    ldr r0, [r1], #4 ; 1 cycle (pipelined with following str)
    str r0, [r2], #4 ; 2 cycles
    ldr r0, [r1], #4 ; 1 cycle (pipelined with following str)
    str r0, [r2], #4 ; 2 cycles
    cmp r1, r3       ; 1 cycle
    blt loop         ; 2 cycles
then we do four copies for every 15 cycles in the optimum case, averaging to 3.75 cycles per word copied, or about 0.9375 cycles per byte copied.

I don't personally want to participate in optimization stuff anymore. I can't take that anymore.
Micro-optimizations like these are almost certainly wasted time, although optimizing memcpy() for ones own hardware might be worth the effort. 

What really bugs me is that the compiler always calls the generic version, even when it knows the pointer alignment and size beforehand.  It could use say memcpy_alignedN_sizeM() for a few different N and M, and just have them as weak symbols that map to memcpy(). Optimizing those might make a difference.

Cramming everything into calling a single function, ignoring all the knowledge about the pointers and size, and then trying to optimize that single function is a bit lunatic, if you think about it carefully.
Title: Re: GCC compiler optimisation
Post by: Kalvin on February 11, 2023, 10:28:10 am
Using DMA for a generic memcpy will complicate things in systems with a preemptive scheduler (as the DMA needs to be shared & synchronized between all tasks in the system), and will make code less portable. I would just keep things simple, and do some manual loop unrolling, resulting simple, portable implementation. For example, manual unrolling by 8 will reduce the loop checking and jump from N to N/8, keeping the pipeline penalty very low.

Code: [Select]
#define MEMCPY_BLOCK_SIZE (8)

void *memcpy_my(void *restrict dest, const void *restrict src, int len)
{
    char *dp = (char *restrict)dest;
    const char *sp = (const char *restrict)src;

    while(len >= MEMCPY_BLOCK_SIZE)
    {
        *dp++ = *sp++; // 1
        *dp++ = *sp++; // 2
        *dp++ = *sp++; // 3
        *dp++ = *sp++; // 4
        *dp++ = *sp++; // 5
        *dp++ = *sp++; // 6
        *dp++ = *sp++; // 7
        *dp++ = *sp++; // 8
        COMPILETIME_ASSERT(MEMCPY_BLOCK_SIZE == 8);
        len -= MEMCPY_BLOCK_SIZE;
    }

    while( len > 0)
    {
        len--;
        *dp++ = *sp++;
    }

    return dest;
}

macro COMPILETIME_ASSERT is just there as a place-holder for the actual compile-time assert, and will warn if the MEMCPY_BLOCK_SIZE and the actual implementation will get out of sync.
Title: Re: GCC compiler optimisation
Post by: peter-h on February 11, 2023, 11:44:58 am
Code: [Select]

        *dp++ = *sp++; // 1
        *dp++ = *sp++; // 2
        *dp++ = *sp++; // 3
        *dp++ = *sp++; // 4
        *dp++ = *sp++; // 5
        *dp++ = *sp++; // 6
        *dp++ = *sp++; // 7
        *dp++ = *sp++; // 8

The problem I found is that for most optimisation levels most of these eight (or four, etc) get removed (and the loop counter is adjusted accordingly). So to be sure you are getting what you want, and the code is proof against future compiler versions, you would have to do it in assembler.

And it is no use saying that if your code breaks with say -O3 then it is broken. Lots of people here have been posting that. This is not true at all if talking to hardware and having to e.g. meet CS timings. The reality of embedded is that these things do matter, and real detail often matters.

Very good point about a DMA version of memcpy needing a mutex. So one would do a DMA version only for a particular thread context. In this case I was trying to optimise the low_level_* functions which interface ETH PHY to LWIP. Some versions of this use zero-copy but that is way too convoluted for me to understand (and lots of people evidently have trouble with it).

Title: Re: GCC compiler optimisation
Post by: Kalvin on February 11, 2023, 12:06:19 pm
The problem I found is that for most optimisation levels most of these eight (or four, etc) get removed (and the loop counter is adjusted accordingly). So you would have to do it in assembler.

Yes, I noticed that also while playing with that particular piece of code in godbolt.org. It is possible to use a suitable #pragma in the source code, or isolate the function into its own compilation unit so that the compiler can then use the best optimization for that particular function. Both solutions will probably require target- and compiler-specific pragmas / compiler options. But it is doable with little effort.

For GCC, the following seems to work with different compiler optimization levels so that the unrolled loop will not get optimized away:
Code: [Select]
#pragma GCC push_options
#pragma GCC optimize "-Os"
#pragma GCC optimize "-fno-tree-loop-distribute-patterns"
void *memcpy_my(void *restrict dest, const void *restrict src, int len)
{
   // body here
}
#pragma GCC pop_options
Title: Re: GCC compiler optimisation
Post by: Siwastaja on February 11, 2023, 02:31:57 pm
Very good point about a DMA version of memcpy needing a mutex.

To circumvent the difficulty of parallel programming, you can always just set up DMA and poll for its completion, doing nothing useful on CPU in the meantime. Wasting the parallel nature of the DMA seems like counterintuitive waste, but if and when DMA is faster than CPU copy, you still save time; even if wasting the possibility of saving even more time.

There can be useful combinations, like start copying aligned words on DMA first, and while that is running, let the CPU copy the few odd bytes.
Title: Re: GCC compiler optimisation
Post by: ataradov on February 11, 2023, 05:56:03 pm
then we do four copies for every 15 cycles in the optimum case, averaging to 3.75 cycles per word copied, or about 0.9375 cycles per byte copied.
2 cycles per branch is optimistic. But otherwise I agree. At the same time nothing stops DMA from doing 4 bytes/cycle in the same setup assuming addresses are aligned. And if addresses are not aligned, then timings for ldr/str are wrong, since they would be converted into multiple aligned accesses internally, so will take longer.

What really bugs me is that the compiler always calls the generic version, even when it knows the pointer alignment and size beforehand.  It could use say memcpy_alignedN_sizeM() for a few different N and M, and just have them as weak symbols that map to memcpy(). Optimizing those might make a difference.
But if you don't mess with compiler setting to much and let it use builtin memcpy(), then it will do almost the same. It does not exactly use different versions, but the version you get quickly checks the alignment and proceeds to do fast transfers if possible.
Title: Re: GCC compiler optimisation
Post by: wek on February 11, 2023, 07:16:07 pm
Some experimental results - 0x1000 word transfers on Disco F4, reset clock settings.

Code: [Select]
                           FIFO   P(src) M(dst)
EXPERIMENT  transfer       thrsh. burst  burst  => t
         0  SRAM1->SRAM2   full     1      1       0x4c2a
         1  SRAM2->SRAM2   full     1      1       0x4c2a
         2  SRAM1->SRAM2   1/4      1      1       0x401e
         3  SRAM2->SRAM2   1/4      1      1       0x401e
         4  SRAM1->SRAM2   full     4      4       0x3819
         5  SRAM1->SRAM2   full     4      1       0x581d
         6  SRAM1->SRAM2   full     1      4       0x4c1d

(I took some old blinky experiment, that's what the timer stuff is there, just ignore it)

JW
Title: Re: GCC compiler optimisation
Post by: ataradov on February 11, 2023, 08:01:01 pm
So, about 4 cycles per transfer.

With 4K words test block you are violating the requirement to not cross 1K boundary for burst transfers. I don't think this would affect timings, but the data is likely to wrap around to 1K boundary, so the code is only useful for time testing. Although with completely aligned blocks bursts would not cross the boundary, so in this case the test is valid even for data.

And I think you might get better performance doing single transfers with 1/2 FIFO threshold. In theory in this scenario the only masters requesting the bus access to SRAM are DMA interfaces, so they would be granted on each cycle. So if you use FIFO as an actual FIFO instead of accumulator for a burst transfer, you should get the state where reading master always has space to place the next value and the writing master always has data in the FIFO. This would not work if DMA does not try to pipeline its requests without bursts though. Which is entirely possible if it is oriented at peripheral transfers.

And just for a teat, I would try to add a block of 8-16 nops in the register poll loop. Just to see if the core generating bus requests somehow affects the result. It should not, but who knows.

Title: Re: GCC compiler optimisation
Post by: wek on February 11, 2023, 08:08:27 pm
> With 4K words test block you are violating the requirement to not cross 1K boundary for burst transfers.

These are 4-beat word bursts, so it's enough to align them at 32-byte boundaries to be sure not to cross the 1k boundary.

> And I think you might get better performance doing single transfers with 1/2 FIFO threshold.

a moment please...

JW

Title: Re: GCC compiler optimisation
Post by: wek on February 11, 2023, 08:26:32 pm
Code: [Select]
                           FIFO   P(src) M(dst)                disturb
EXPERIMENT  transfer       thrsh. burst  burst  => t           1=SRAM2wr 2=SRAM1wr 3=SRAM1rd
         0  SRAM1->SRAM2   full     1      1       0x4c2a      0x502b    0x502b    0x502b
         1  SRAM2->SRAM2   full     1      1       0x4c2a      0x502b    0x4c25    0x4c25
         2  SRAM1->SRAM2   1/4      1      1       0x401e      0x4023
         3  SRAM2->SRAM2   1/4      1      1       0x401e      0x4023
         4  SRAM1->SRAM2   full     4      4       0x3819      0x4016    0x401a    0x401a
         5  SRAM1->SRAM2   full     4      1       0x581d      0x601a    0x601a
         6  SRAM1->SRAM2   full     1      4       0x4c1d      0x4c1c
         7  SRAM1->SRAM2   1/2      1      1       0x4024
----
SW loop   => t2 = 0x6005

  for (uint32_t i = 0; i < 0x1000; i++) {
    *(volatile uint32_t *)0x2001C000 = *(volatile uint32_t *)0x20001000;
  }
 80002ac: f44f 5280 mov.w r2, #4096 ; 0x1000
 80002b0: 6829      ldr r1, [r5, #0]
 80002b2: 6021      str r1, [r4, #0]
 80002b4: 3a01      subs r2, #1
 80002b6: d1fb      bne.n 80002b0 <main+0xdc>

Instead of nops in the DMA-done-wait loop, I added read/write to the SRAMs. In cases, where it made no difference, tried to add/remove a couple (literally) of NOPs but that didn't make any difference either.

One of the things which is sort of a mystery to me is the non-integer number of cycles per transfer.

The software loop is 6-cycle.

JW
Title: Re: GCC compiler optimisation
Post by: ataradov on February 11, 2023, 08:36:17 pm
The only explanation I have for this is that DMA does not pipeline its bus access. It waits for the completion of the transfer and then starts a new one. In this case maximum length bust would give the best performance, since at least data phase of the transfer would be pipelined. This is not unexpected from a general purpose DMA, I guess. Most of the time it works with peripherals where pipelined access makes no sense.

As far as non-integer number of cycles, I would exclude the register setup from measurement. Just measure the time from channel enable to the loop end. Different constants may generate slightly different versions of the code, which changes alignment.  And may be try different transfer size to see how it scales per transfer overhead vs fixed overhead.

And the code one is unfair, since it is just a word transfer with no increments. I don't expect write-back versions of the instructions to be slower, but who knows. At the very least, ldr/str instructions would be 32-bit instructions, which might matter for the code size and it fitting into the flash accelerator buffers, which would be an issue at real clock speeds.
Title: Re: GCC compiler optimisation
Post by: peter-h on February 11, 2023, 09:19:07 pm
So DMA is roughly 3/2 i.e. 1.5x faster. Very useful to know!

The application area where this is really important is going to be very narrow, especially given that it is a single-thread proposition.

In my product I have moved all SPI transfers to DMA. Even the very slow ones, for uniformity. Runs very nicely. But all the SPI peripherals are mutex protected for thread safety, and the SPI is auto re-initialised (inside the mutex region, obviously) when peripherals are switched. Took a while to get this to work right. One could do the same with mem-mem DMA. But then my SPI can run up to 21MHz clock and - some past threads - software can't keep up with that, largely because peripherals run very slowly so their register access is slow.
Title: Re: GCC compiler optimisation
Post by: wek on February 11, 2023, 10:32:50 pm
> The only explanation I have for this [...]

I'm afraid we don't have enough information to make definitive conclusions - and, more importantly, predictions. What would be impact of changing transfer width at either side? What's the impact of slow memories (FLASH,  external)? What's the impact of other streams of the same DMA running concurrently? All this could be benchmarked but the data would be overwhelming and even harder to interpret.

Of course, it would be far better if an insider would explain the exact working of this DMA clearly and concisely; but ST is not willing to engage in anything beyond clicking in CubeMX. I digress.

> As far as non-integer number of cycles, I would exclude the register setup from measurement

That's quite obviously the cca 0x20 portion of the timing, the true timing would be most probably 0xXY00. But still, Y is nonzero, and we are talking 0x1000 transfers (e.g. in first case, 0x4c00 cycles per 0x1000 transfers, so its 4 and 3/4 of cycle per transfer, or, 3 transfers out of 4 take 5 cycle and the fourth 4 cycle).

> And the code one is unfair

True.

I couldn't coax gcc into doing what I wanted, it kept incrementing only one of the pointers and adding a constant to it to get the other, this resulted in 8 cycles per transfer. I concocted this (please don't laugh too loudly, I am not expert at gcc inline asm):
Code: [Select]
    uint32_t a1 = 0, a2 = 0, a3 = 0, a4 = 0;

    __asm__ volatile (
       "ldr  %[src], =0x20001000"  "\n\t"
       "ldr  %[dst], =0x2001C000"  "\n\t"
       "mov  %[cnt], 0x1000"       "\n"
"1:\t" "ldr  %[tmp], [%[src]], #4" "\n\t"
       "str  %[tmp], [%[dst]], #4" "\n\t"
       "subs %[cnt], #1"           "\n\t"
       "bne  1b"
       : [src] "=r" (a1)
       , [dst] "=r" (a2)
       , [cnt] "=r" (a3)
       , [tmp] "=r" (a4)
    );
and this yields 7 cycles per transfer. More precisely 0x7006, yes some cycles for setup etc.

This runs at the default 16MHz clock so no waitstates. Without going through the hassle of firing up the crystal oscillator and PLL (which is irrelevant for timing), I just set the latencies and enable the jumpcache a.k.a. "ART accelerator"
Code: [Select]

#if (1) // just mimic high gear by setting latency
  FLASH->ACR =  0
    | FLASH_ACR_ICEN          // meantime, configure flash for maximum performance - enable both caches
    | FLASH_ACR_DCEN
    | FLASH_ACR_PRFTEN        // and enable prefetch
    | FLASH_ACR_LATENCY_5WS   // for VCC>2.7V and FSYS>150MHz, 5 waitstate is appropriate
  ;
  RCC->CFGR = (RCC->CFGR & (~(0
    | RCC_CFGR_PPRE1
    | RCC_CFGR_PPRE2
  ))) | (0
    | RCC_CFGR_PPRE1_DIV4   // APB1 prescaler set to 4 -> APB1 clock = 45MHz (required to be <= 45MHz)
    | RCC_CFGR_PPRE2_DIV2   // APB2 prescaler set to 2 -> APB2 clock = 90MHz (required to be <= 90MHz)
  );
#endif
and... the result is 0x7016, so the literal fetches and the first jump suffered the latency, the remaining 0xFFF jumps were served from the jumpcache; plus maybe some of the setup code exhausted the prefetch or something similar.

JW
Title: Re: GCC compiler optimisation
Post by: abyrvalg on February 12, 2023, 11:07:11 am
Assembly implementations could utilize LDM/STM instructions to move more data per loop iteration. Some common library memcpy() (ARMCC’s? Not sure which one) chooses between 3 loops based on data size/alignment - 16/4/1.
Title: Re: GCC compiler optimisation
Post by: peter-h on February 12, 2023, 01:33:22 pm
An asm version of the memcpy I posted above (which moves 32 bits at a time until the last few (or 0) residual bytes) would be very useful. I have done too little arm32 asm to work it out myself though.

On the wider topic I am surprised that Cube doesn't come with this stuff. After all, you specify the CPU in a config file, and that is used to throw in different libs and all kinds of stuff.

Predictably, this was discussed before
https://www.eevblog.com/forum/programming/should-memcmp-be-faster-than-a-loop/ (https://www.eevblog.com/forum/programming/should-memcmp-be-faster-than-a-loop/)

This is the asm from my above C function, with -Og

Code: [Select]
         
memcpy_fast:
0804ff4c:   cmp     r2, #3
0804ff4e:   bhi.n   0x804ff6c <memcpy_fast+32>
0804ff50:   mov     r3, r0
0804ff52:   b.n     0x804ff8e <memcpy_fast+66>
0804ff54:   ldrb.w  r2, [r1], #1
0804ff58:   strb.w  r2, [r3], #1
0804ff5c:   mov     r2, r12
0804ff5e:   add.w   r12, r2, #4294967295
0804ff62:   cmp     r2, #0
0804ff64:   bne.n   0x804ff54 <memcpy_fast+8>
0804ff66:   ldr.w   r4, [sp], #4
0804ff6a:   bx      lr
0804ff6c:   mov     r3, r0
0804ff6e:   cmp     r2, #3
0804ff70:   bls.n   0x804ff8e <memcpy_fast+66>
0804ff72:   push    {r4}
0804ff74:   ldr.w   r4, [r1], #4
0804ff78:   str.w   r4, [r3], #4
0804ff7c:   subs    r2, #4
0804ff7e:   cmp     r2, #3
0804ff80:   bhi.n   0x804ff74 <memcpy_fast+40>
0804ff82:   b.n     0x804ff5e <memcpy_fast+18>
0804ff84:   ldrb.w  r2, [r1], #1
0804ff88:   strb.w  r2, [r3], #1
0804ff8c:   mov     r2, r12
0804ff8e:   add.w   r12, r2, #4294967295
0804ff92:   cmp     r2, #0
0804ff94:   bne.n   0x804ff84 <memcpy_fast+56>
0804ff96:   bx      lr

With -O0 you get this

Code: [Select]
          memcpy_fast:
0804ff4c:   cmp     r2, #3
0804ff4e:   bhi.n   0x804ff6c <memcpy_fast+32>
395        char *dst = dst0;
0804ff50:   mov     r3, r0
0804ff52:   b.n     0x804ff8e <memcpy_fast+66>
420        *dst++ = *src++;
0804ff54:   ldrb.w  r2, [r1], #1
0804ff58:   strb.w  r2, [r3], #1
419        while (len0--)
0804ff5c:   mov     r2, r12
0804ff5e:   add.w   r12, r2, #4294967295
0804ff62:   cmp     r2, #0
0804ff64:   bne.n   0x804ff54 <memcpy_fast+8>
424       }
0804ff66:   ldr.w   r4, [sp], #4
0804ff6a:   bx      lr
404        aligned_dst = (uint32_t*)dst;
0804ff6c:   mov     r3, r0
407        while (len0 >= 4)
0804ff6e:   cmp     r2, #3
0804ff70:   bls.n   0x804ff8e <memcpy_fast+66>
393       {
0804ff72:   push    {r4}
409        *aligned_dst++ = *aligned_src++;
0804ff74:   ldr.w   r4, [r1], #4
0804ff78:   str.w   r4, [r3], #4
410        len0 -= 4;
0804ff7c:   subs    r2, #4
407        while (len0 >= 4)
0804ff7e:   cmp     r2, #3
0804ff80:   bhi.n   0x804ff74 <memcpy_fast+40>
0804ff82:   b.n     0x804ff5e <memcpy_fast+18>
420        *dst++ = *src++;
0804ff84:   ldrb.w  r2, [r1], #1
0804ff88:   strb.w  r2, [r3], #1
419        while (len0--)
0804ff8c:   mov     r2, r12
0804ff8e:   add.w   r12, r2, #4294967295
0804ff92:   cmp     r2, #0
0804ff94:   bne.n   0x804ff84 <memcpy_fast+56>
0804ff96:   bx      lr

I've done megabytes of asm but struggle with the weird ways of the GCC compiler. Obviously stuff like this

Code: [Select]
aligned_dst = (uint32_t*)dst;
aligned_src = (uint32_t*)src;

is just for the benefit of C and if writing in asm you hardly consider it; a register value can be used as an address directly.
Title: Re: GCC compiler optimisation
Post by: abyrvalg on February 12, 2023, 03:04:32 pm
The library memcpy mentioned earlier:
Code: [Select]
__rt_memcpy                            ; unaligned version
                CMP             R2, #3
                BLS             loc_59F00
                ANDS            R12, R0, #3
                BEQ             loc_59ED2
                LDRB            R3, [R1],#1
                CMP             R12, #2
                ADD             R2, R12
                IT LS
                LDRBLS          R12, [R1],#1
                STRB            R3, [R0],#1
                IT CC
                LDRBCC          R3, [R1],#1
                SUB             R2, R2, #4
                IT LS
                STRBLS          R12, [R0],#1
                IT CC
                STRBCC          R3, [R0],#1

loc_59ED2                               ; CODE XREF: __rt_memcpy+A↑j
                ANDS            R3, R1, #3
                BEQ             __aeabi_memcpy8
                SUBS            R2, #8

loc_59EDC                               ; CODE XREF: __rt_memcpy+54↓j
                BCC             loc_59EF0
                LDR             R3, [R1],#4
                SUBS            R2, #8
                LDR             R12, [R1],#4
                STM             R0!, {R3,R12}
                B               loc_59EDC
; ---------------------------------------------------------------------------

loc_59EF0                               ; CODE XREF: __rt_memcpy:loc_59EDC↑j
                ADDS            R2, R2, #4
                ITT PL
                LDRPL           R3, [R1],#4
                STRPL           R3, [R0],#4
                NOP 

loc_59F00                               ; CODE XREF: __rt_memcpy+2↑j
                LSLS            R2, R2, #0x1F
                ITT CS
                LDRBCS          R3, [R1],#1
                LDRBCS          R12, [R1],#1
                IT MI
                LDRBMI          R2, [R1],#1
                ITT CS
                STRBCS          R3, [R0],#1
                STRBCS          R12, [R0],#1
                IT MI
                STRBMI          R2, [R0],#1
                BX              LR
; End of function __rt_memcpy

__aeabi_memcpy8          ; 8-byte aligned version       
                PUSH            {R4,LR}
                SUBS            R2, #0x20
                BCC             loc_59FC6

loc_59FB0                               ; CODE XREF: __aeabi_memcpy8+1A↓j
                ; 32 bytes per iteration loop
                LDM             R1!, {R3,R4,R12,LR} 
                SUBS            R2, #0x20
                STM             R0!, {R3,R4,R12,LR}
                LDM             R1!, {R3,R4,R12,LR}
                STM             R0!, {R3,R4,R12,LR}
                BCS             loc_59FB0

loc_59FC6                               ; CODE XREF: __aeabi_memcpy8+4↑j
                MOVS            R12, R2,LSL#28
                ; 16-byte remainder
                ITT CS
                LDMCS           R1!, {R3,R4,R12,LR}
                STMCS           R0!, {R3,R4,R12,LR}
                ; 8-byte remainder
                ITT MI
                LDMMI           R1!, {R3,R4}
                STMMI           R0!, {R3,R4}
                POP             {R4,LR}
                MOVS            R12, R2,LSL#30
                ; 4-byte remainder
                ITT CS
                LDRCS           R3, [R1],#4
                STRCS           R3, [R0],#4
                IT EQ
                BXEQ            LR
                LSLS            R2, R2, #0x1F
                ; 2-byte remainder
                IT CS
                LDRHCS          R3, [R1],#2
                ; 1-byte remainder
                IT MI
                LDRBMI          R2, [R1],#1
                IT CS
                STRHCS          R3, [R0],#2
                IT MI
                STRBMI          R2, [R0],#1
                BX              LR
; End of function __aeabi_memcpy8

The compiler emits direct __aeabi_memcpy8 calls when it knows the data alignment and __rt_memcpy in all other cases.
Title: Re: GCC compiler optimisation
Post by: cv007 on February 12, 2023, 03:40:47 pm
Simple question - why do you have to move big chunks of memory-memory rather than use something like a description block to prevent the need to do so?
Title: Re: GCC compiler optimisation
Post by: peter-h on February 12, 2023, 04:57:53 pm
Because I am moving between two "blocks" which I don't understand and which almost nobody else understands.

There is in this case a zero-copy approach but despite having been banging about for ~10 years is complicated and I have never seen code known to be working and robust (for the 32F4).
Title: Re: GCC compiler optimisation
Post by: Siwastaja on February 12, 2023, 06:52:08 pm
Yeah, zero-copy solutions are nice when you can properly engineer the whole thing. In practice, one needs to just memcpy things, but OTOH, thankfully, people also tend to overestimate the cost of doing so.

Zero-copy has not only the advantage of saving time; it also saves memory. All kind of stupid buffers can take a lot of RAM in embedded system, when every layer wants to have their own buffers.
Title: Re: GCC compiler optimisation
Post by: peter-h on February 13, 2023, 01:57:35 pm
Someone pointed out that a memcpy implemented as 4 byte moves will fail if the gap between the two buffers is less than 4 bytes, which is why, supposedly, memcpy is normally done in byte mode.

Can anybody relate to this? I don't get it at all.

Normal memcpy doesn't support overlapping blocks. For that you use memmove.

Code: [Select]
void * memcpy_fast (void *__restrict dst0, const void *__restrict src0, size_t len0)
{

char *dst = dst0;
const char *src = src0;
uint32_t *aligned_dst;
const uint32_t *aligned_src;

// If the size is >=4 then do 32 bit moves, until exhausted

if ( len0 >=4 )
{
aligned_dst = (uint32_t*)dst;
aligned_src = (uint32_t*)src;

while (len0 >= 4)
{
*aligned_dst++ = *aligned_src++;
len0 -= 4;
}

dst = (char*)aligned_dst;
src = (char*)aligned_src;
}

// Finish with any single byte moves

while (len0--)
*dst++ = *src++;

return dst0;

}
Title: Re: GCC compiler optimisation
Post by: eutectique on February 13, 2023, 02:15:12 pm
Suppose two memory blocks, 0x10 bytes in size, src starts at address 0x100, dst starts at address 0x108 -- that's the definition of overlapping blocks.

memcpy() that uses src++ and dst++ will fail, byte-wide or word-wide, doesn't matter.
Title: Re: GCC compiler optimisation
Post by: NorthGuy on February 13, 2023, 04:31:40 pm
Someone pointed out that a memcpy implemented as 4 byte moves will fail if the gap between the two buffers is less than 4 bytes

If you move in consecutive blocks starting from the lower end, then regardless of the bock size, memcpy wil fail if both of these conditions are met:

Code: [Select]
src < dst < (src+len)
len > block_size

and succeed otherwise.
Title: Re: GCC compiler optimisation
Post by: ejeffrey on February 13, 2023, 05:29:47 pm
Someone pointed out that a memcpy implemented as 4 byte moves will fail if the gap between the two buffers is less than 4 bytes, which is why, supposedly, memcpy is normally done in byte mode.

Can anybody relate to this? I don't get it at all.

I don't get it either.  A properly working memcpy implementation should neither read bytes outside the source buffer nor write bytes outside the destination.  If this requirement is satisfied, it should not fail as long as the buffers don't overlap (which as you say is not supported, and requires memmove).  Some implementations do allow reading past the end of the source, as long as no architectural boundaries are crossed (i.e., the over-read can't cause a fault), typically to the end of an aligned word.  But since the extra bytes won't be written, it doesn't matter whether they are new or old data.

If the source and destination buffers are disjoint, but share the same 4 byte word, then they can't both be 4 byte aligned.  The reason simple memcpy implementations work in byte mode is entirely for alignment purposes.  memcpy doesn't place any alignment requirements on the buffers or the sizes.  So memcpy implementations that want to use larger word sizes need to handle buffers that are not aligned, lengths that are not a multiple of the word size, different alignment between source and destination, and do so without causing faults or severe misaligned access penalties.

Optimized memcpy implementations obviously *do* check the sizes and alignments and try to use larger chunks when possible.

Title: Re: GCC compiler optimisation
Post by: peter-h on February 13, 2023, 07:19:03 pm
The code I posted above uses 32 bit moves until there are 0-3 bytes left, so it can't be overrunning either buffer.
Title: Re: GCC compiler optimisation
Post by: eutectique on February 13, 2023, 08:24:54 pm
The code I posted above uses 32 bit moves until there are 0-3 bytes left, so it can't be overrunning either buffer.

Try this:

Code: [Select]
src [1][2][3]
dst       [?][?][?]
Title: Re: GCC compiler optimisation
Post by: eutectique on February 13, 2023, 08:53:29 pm
Or better this:

Code: [Select]
src [1][2][3][4][5]
dst             [ ][ ][ ][ ][ ]
Title: Re: GCC compiler optimisation
Post by: peter-h on February 13, 2023, 09:20:27 pm
Those are overlapping buffers.

I wrote

Quote
Someone pointed out that a memcpy implemented as 4 byte moves will fail if the gap between the two buffers is less than 4 bytes
Title: Re: GCC compiler optimisation
Post by: brucehoult on February 14, 2023, 08:55:52 am
Those are overlapping buffers.

I wrote

Quote
Someone pointed out that a memcpy implemented as 4 byte moves will fail if the gap between the two buffers is less than 4 bytes

That's not failing. The memcpy() spec says the results are undefined if src and dst overlap.

If you give memcpy() overlapping buffers then it is you that has failed to honour your side of the contract, not memcpy().
Title: Re: GCC compiler optimisation
Post by: peter-h on February 14, 2023, 09:04:32 am
OK we got wires crossed.

The issue I described above (buffers with a gap between them of < 4 bytes) doesn't exist AFAICT.
Title: Re: GCC compiler optimisation
Post by: peter-h on May 25, 2023, 08:30:32 pm
I've just watched this video

https://www.youtube.com/watch?v=w0sz5WbS5AM (https://www.youtube.com/watch?v=w0sz5WbS5AM)

and it convinces me more than ever that a huge amount of work has gone into optimisations which yield no useful runtime benefits whatsoever. For example how time critical can swapping the four bytes of a uint32_t possibly be?
Title: Re: GCC compiler optimisation
Post by: voltsandjolts on May 25, 2023, 08:51:30 pm
how time critical can swapping the four bytes of a uint32_t possibly be?

If you need to do it 1e6 times per second, then pretty darn critical.
Title: Re: GCC compiler optimisation
Post by: peter-h on May 25, 2023, 09:18:13 pm
There is a lot of coding style dependency involved in triggering the pattern recognitions.
Title: Re: GCC compiler optimisation
Post by: ataradov on May 25, 2023, 09:28:27 pm
There is a lot of coding style dependency involved in triggering the pattern recognitions.
So? Hardware design tools are all based on very strict pattern for the tool to recognize your intent and generate optimized design. And hardware people deal with that.

If you are doing embedded stuff where nothing matters, then you may not appreciate all that optimization.  It does not mean it is useless.

Networking code involves a lot of swaps and recognizing that pattern may lead to a much faster code. Recognizing code that could be automatically vectored, results in huge performance increase on platforms that have vector hardware.

Feel free to use TCC or something like that if you don't want optimizations and don't like them even being in the compiler .
Title: Re: GCC compiler optimisation
Post by: peter-h on May 25, 2023, 09:34:37 pm
Quote
feel free to use TCC or something like that if you don't want optimizations and don't like them even being in the compiler .

WOW.
Title: Re: GCC compiler optimisation
Post by: hans on May 26, 2023, 12:05:27 pm
What is easier to understand:
Quote
const int wrap = 8;
int idx = x % wrap;
Or:
Quote
const int wrap = 8;
static_assert((wrap & (wrap - 1)) == 0, "Wrap is not a power of 2");
int idx = x & (wrap - 1);
Maybe we as embedded programmers feel more comfortable with the 2nd one, but that's pretty domain specific knowledge. The average programmer will understand the first better, however, if a compiler is able to make the step to implementation #2 as long as wrap \in 2^N (which is put down as an implementation constraint of #2), then that's great. After all, almost no compiler is going to naively compile the modulo using a REM instruction (integer division unit) as its super slow, so I'm happy that compilers are smarter than me to find bit and integer manipulation tricks. Some of these tricks, like the whitespace check function, can also work on any platform that has a barrel shifter. These kind of optimizations can benefit x86, RISC-V, ARM and MIPS all in 1 go.

Vector code is all the hype these days. Unfortunately GCC doesn't generate SIMD instructions for Cortex-M4 cores, yet. These instructions are a must for DSP applications. The other day I optimized an algorithm that took 2M+ cycles/second back down to only a few hunderd kcyc/s. It would be amazing if the compiler could do this by itself - as I've to check between a naive MATLAB implementation (including several Fourier convolutions and complex number manipulations) and its SIMD optimized version. The problem with the SIMD version is that I can't easily unit test it.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on May 26, 2023, 07:20:32 pm
No doubt that compiler optimizations are a big plus in a whole range of case.

The ability to map to specific target instructions of course all depends on the target and current state of the compiler. GCC currently does a better job for x86, not really surprising.

Since some talked about byte swapping here goes:

Code: [Select]
#include <stdint.h>

uint32_t ByteSwap32( uint32_t val )
{
    val = ((val << 8) & 0xFF00FF00 ) | ((val >> 8) & 0xFF00FF );
    return (val << 16) | (val >> 16);
}

gcc -O2:
Code: [Select]
ByteSwap32(unsigned int):
        mov     eax, edi
        bswap   eax
        ret

gcc -O1:
Code: [Select]
ByteSwap32(unsigned int):
        mov     eax, edi
        sal     eax, 8
        and     eax, -16711936
        shr     edi, 8
        and     edi, 16711935
        or      eax, edi
        rol     eax, 16
        ret

gcc -O0:
Code: [Select]
ByteSwap32(unsigned int):
        push    rbp
        mov     rbp, rsp
        mov     DWORD PTR [rbp-4], edi
        mov     eax, DWORD PTR [rbp-4]
        sal     eax, 8
        and     eax, -16711936
        mov     edx, eax
        mov     eax, DWORD PTR [rbp-4]
        shr     eax, 8
        and     eax, 16711935
        or      eax, edx
        mov     DWORD PTR [rbp-4], eax
        mov     eax, DWORD PTR [rbp-4]
        rol     eax, 16
        pop     rbp
        ret

For ARM (32), not too bad either:

gcc -O2:
Code: [Select]
ByteSwap32(unsigned int):
        rev     r0, r0
        bx      lr

gcc -O1:
Code: [Select]
ByteSwap32(unsigned int):
        lsls    r3, r0, #8
        and     r3, r3, #-16711936
        lsrs    r0, r0, #8
        and     r0, r0, #16711935
        orrs    r0, r0, r3
        ror     r0, r0, #16
        bx      lr

Which is faster?

While it was maybe a bit tongue-in-cheek, TCC is actually not that bad if you don't care about optimizations. It's ultra-fast to compile and produces correct code as far as I tested it.

Now of course you can tell the difference in terms of optimization.

Code: [Select]
00000000 <ByteSwap32>:
   0:   e1a0c00d        mov     ip, sp
   4:   e92d0003        push    {r0, r1}
   8:   e92d5800        push    {fp, ip, lr}
   c:   e1a0b00d        mov     fp, sp
  10:   e1a00000        nop                     ; (mov r0, r0)
  14:   e59b000c        ldr     r0, [fp, #12]
  18:   e1a00400        lsl     r0, r0, #8
  1c:   e59f1000        ldr     r1, [pc]        ; 24 <ByteSwap32+0x24>
  20:   ea000000        b       28 <ByteSwap32+0x28>
  24:   ff00ff00                        ; <UNDEFINED> instruction: 0xff00ff00
  28:   e0000001        and     r0, r0, r1
  2c:   e59b100c        ldr     r1, [fp, #12]
  30:   e1a01421        lsr     r1, r1, #8
  34:   e59f2000        ldr     r2, [pc]        ; 3c <ByteSwap32+0x3c>
  38:   ea000000        b       40 <ByteSwap32+0x40>
  3c:   00ff00ff        ldrshteq        r0, [pc], #15
  40:   e0011002        and     r1, r1, r2
  44:   e1800001        orr     r0, r0, r1
  48:   e58b000c        str     r0, [fp, #12]
  4c:   e59b000c        ldr     r0, [fp, #12]
  50:   e1a00800        lsl     r0, r0, #16
  54:   e59b100c        ldr     r1, [fp, #12]
  58:   e1a01821        lsr     r1, r1, #16
  5c:   e1800001        orr     r0, r0, r1
  60:   e89ba800        ldm     fp, {fp, sp, pc}
;D
Title: Re: GCC compiler optimisation
Post by: peter-h on May 26, 2023, 08:12:12 pm
The various GCC levels just trim the code more and more, so why not?

Your last example is what the H8/300 GNU compiler produced in 1995 :) I could have used that for a customer programmable product back then but it was awful. So I chose the Hitech one which was vastly better but had to be sold for £450.

My earlier comments were on a different thing.
Title: Re: GCC compiler optimisation
Post by: ataradov on May 26, 2023, 08:34:02 pm
To be clear, I did not flame TCC, it is an excellent compiler. It is just a good example of total no-optimization compiler. It is something you don't want at all for general purpose programming.
Title: Re: GCC compiler optimisation
Post by: SiliconWizard on May 26, 2023, 08:42:05 pm
To be clear, I did not flame TCC, it is an excellent compiler. It is just a good example of total no-optimization compiler. It is something you don't want at all for general purpose programming.

Yep. Well, TCC tends to do a bit worse than GCC -O0, but not a whole lot worse. So, for someone that would consistently compile their code at -O0, TCC would certainly be an alternative.

(The only thing is that it would take a bit of work to use in your toolchain compared to using a vendor-supplied GCC with everything already set up.)
Title: Re: GCC compiler optimisation
Post by: bson on May 26, 2023, 09:25:37 pm
Why bother with this "systick" interrupt when you're at a low level stage?
Just use the DWT cycle counter, as is often suggested in this forum. It just needs to be enabled. I use this for enabling it in C:
Code: [Select]
void DWT_Init(void)
{
if (! (CoreDebug->DEMCR & CoreDebug_DEMCR_TRCENA_Msk))
CoreDebug->DEMCR |= CoreDebug_DEMCR_TRCENA_Msk;

DWT->CYCCNT = 0;
DWT->CTRL |= DWT_CTRL_CYCCNTENA_Msk;
}

Enabling it in assembly is also just a couple instructions if you need to do this in the startup code or something.

Then you can read the DWT->CYCCNT register. Resolution is the system clock period. Doesn't require any interrupt or any peripheral.
Why check DMCR & TRCENA before setting TRCENA?  Is there something in the STM32 (F) implementation of the DWT or ITM modules that necessitates this?
I've never checked it and haven't noticed any problems, though I use ITM (for SWO) and not really DWT.  The DWT though has useful optimization data so I may start using it!

Title: Re: GCC compiler optimisation
Post by: brucehoult on May 27, 2023, 01:58:10 am
Since some talked about byte swapping here goes:

riscv32 if you include "_zbb" in the arch string. gcc needs -O2 or -Os, with clang just -O1 is enough:

Code: [Select]
ByteSwap32:
        rev8    a0,a0
        ret
Title: Re: GCC compiler optimisation
Post by: NorthGuy on May 27, 2023, 01:50:15 pm
Since some talked about byte swapping here goes:

These kind of optimizations are needed because C doesn't have operations for byte swaps, rotations, bit settings etc. Therefore, the compiler needs to recognize common patterns which you would use for these and then emit simple commands.
Title: Re: GCC compiler optimisation
Post by: peter-h on May 27, 2023, 06:00:57 pm
Quote
To be clear,

To be clear

Quote
Therefore, the compiler needs to recognize common patterns

that's exactly the point I was trying to make. This code replacement stuff is coding style dependent.

And if your code depends on the speed, then you need to do a lot of regression testing (especially timing stuff with a scope / logic analyser) after each compiler version change. And examine the asm generated to make sure it isn't working by accident.

And if your code doesn't depend on the speed, well, then the time invested by the compiler writer was a bit pointless ;) Delivering a human-readable linkfile syntax would save somebody a lot more time, for example.

If I was writing code which actually needs the speed I would do an asm function for it.

Title: Re: GCC compiler optimisation
Post by: ataradov on May 27, 2023, 06:17:51 pm
And if your code depends on the speed, then you need to do a lot of regression testing (especially timing stuff with a scope / logic analyser) after each compiler version change.
Real projects do that anyway. This is why nobody changes the versions of the tools mid projects. And version change on a big project usually takes weeks of evaluation.

For small projects you can do whatever you want, compiler makers don't target you.

And if your code doesn't depend on the speed, well, then the time invested by the compiler writer was a bit pointless
even if something does not break if it is a bit slower, it does not mean that it does not benefit from the added performance.
Title: Re: GCC compiler optimisation
Post by: AVI-crak on May 27, 2023, 08:44:22 pm
One very crooked company made usb hardware buffers with holes (like cheese). Real 16bit data, 16bit empty, 16bit data, 16bit empty, and so on. To transfer data, the company wrote a simple code that GCC optimized to 32bit reads. GCC knows that the buffer is always clear before starting work, and therefore performs the addition with a shift (it is more convenient for it).
You can look for a data transfer error indefinitely when you accidentally deleted a buffer flush.
So, DMA can read sparse data, and write without emptiness. Able to do in the opposite direction (creates cheese). But the configurator from the company prohibits such DMA settings.
Title: Re: GCC compiler optimisation
Post by: brucehoult on May 28, 2023, 12:34:00 am
If I was writing code which actually needs the speed I would do an asm function for it.

It's safer.

Other than needing 3 or 4 or 5 different asm versions if the code is going to be widely used. Always keep a plain C (with idioms you hope will be recognised) as fallback when the code is compiled for an ISA you don't expect.
Title: Re: GCC compiler optimisation
Post by: westfw on May 28, 2023, 02:06:07 am
A compiler that optimizes a byteswap of memory into a load (at memory speed - 10s of nanoseconds) followed by a SWAP instruction (sub-nanosecond speed, once the value is in a register) can handle arbitrary network traffic in any endianness with essentially no performance penalties.  Allowing Intel and AMD to sell into significant markets that used to be dominated by MIPS CPUs (which could be configured for different endianness.)
As of a dozen year ago, that's making switching decisions (in software) for over 4 million packets/s.