EEVblog® Electronics Community Forum
Electronics => Microcontrollers => Topic started by: microvia on August 23, 2026, 11:18:36 am
-
This is only a warning after discovering that when using optimization level 3 (-O3), the XC8 v4.00 compiler may clear some RAM banks unexpectedly when a PIC16F1825 interrupt routine is about to execute.
The reason was found by two talented programmers on the Microchip forum, who I thank for their time and help. I don't know how many MCUs are concerned.
Those interested may check the initial thread at https://forum.microchip.com/s/topic/a5CV4000001HN2rMAG/t406124 (https://forum.microchip.com/s/topic/a5CV4000001HN2rMAG/t406124).
The detailed explanation can be found from Post #32 at https://forum.microchip.com/s/topic/a5CV4000001HN2rMAG/t406124?comment=P-3014555 (https://forum.microchip.com/s/topic/a5CV4000001HN2rMAG/t406124?comment=P-3014555).
I personally decided to stay with optimization level 2 until MicroChip investigates a ticket that I opened for this purpose.
-
Microchip has reproduced the bug and is currently investigating.
It has been found that the bug not only clears all RAM at the vector interrupt, but also wastes 27% of the code with useless instructions.
Those who want to use optimization level 3 should download and install XC8 v3.00 from the Microchip archive (https://www.microchip.com/en-us/tools-resources/archives/mplab-ecosystem (https://www.microchip.com/en-us/tools-resources/archives/mplab-ecosystem)), and then download the free license key by clicking Access Legacy Compiler License (https://www.microchip.com/en-us/tools-resources/develop/mplab-xc-compilers/xc8 (https://www.microchip.com/en-us/tools-resources/develop/mplab-xc-compilers/xc8)) and execute the BATch file.
-
After investigating this over the past few days this optimizer bug appears to affect the mid-range and enhanced PIC16F controllers. PIC10F and PIC12F may also be affected.
A strange set of statements will provoked the XC8 compiler to manifest this bug. I am not surprised that the Microchip Q/A testers did not find this. The Original Poster shared an example that used an ISR and some bit manipulation of a long data type variables. There are likely other ways, but this is the example that reliably reproduces this fault.
-
I am not surprised that the Microchip Q/A testers did not find this.
Do they lack automated testing facilities?
Many compilers tent to produce some nonsense code that's difficult to debug at high optimization levels.
With GCC/Intel C++/x86 compilers on PC it's a common thing to not go beyond -O2 as things may break.
-
Do they lack automated testing facilities?
It would seem that Microchip must use automated tests as this appears required to get through MISRA Compliance.
This specific bug behavior needs a crazy quilt of statements and maximum optimizations to provoke. It would take a test engineer of extraordinarily malicious intent to create test vectors to find it. This is why Microchip depends on users to find stuff like this for free.
-
With GCC/Intel C++/x86 compilers on PC it's a common thing to not go beyond -O2 as things may break.
Depends on what you mean by common. Actual compiler bugs are rare (sure, when they exist, they are somewhat more likely to trigger at higher optimization levels). Broken code which randomly cures and breaks again depending on optimization level is more common, but even that, nah, not that common, only some people buy the rationalization story ("I don't need to fix my code if I blame the compiler, and blaming others is easier for me than fixing my code" - I wish we got thinking summaries out of unaligned humans, like we get out of unaligned AI!)
But yeah, bugs happen, and of course any serious software project would have testing, but compilers are very complicated beasts so some bugs go past testing, that's life and it happens in every software project larger than "Hello World".
Good first assumption is always suspect your own code, and as the last thing, compiler bug, while not totally ruling out that possibility.
Routinely compiling with different optimization levels is beneficial if it surfaces bugs in your own code. You can then investigate and fix those bugs. Finding a compiler bug instead is so rare it's not a good reason not to use all available optimization levels, IMHO.
-
Good first assumption is always suspect your own code, and as the last thing, compiler bug, while not totally ruling out that possibility.
Well in case if something works totally correct with -O2, but breaks with -O3, or even worse, works OK with -O1, but breaks with -O2, whom fault is that?
-
Good first assumption is always suspect your own code, and as the last thing, compiler bug, while not totally ruling out that possibility.
Well in case if something works totally correct with -O2, but breaks with -O3, or even worse, works OK with -O1, but breaks with -O2, whom fault is that?
Almost always it turns out to be a bug in the program. If it breaks and "unbreaks" by switching between -O2 and -O3, that's great - you have a reproducible way to make the bug appear and disappear, and can compare, investigate.
Timing and race conditions are the most typical such bugs. For example, using a shared variable within multiple threads of execution without proper mutexing/locking. Changing the optimization setting changes the timing relationship, just like adding random delays in the code does, so the bug can appear and disappear depending on the setting. The only robust fix is to find and fix the bug itself, not to find the compiler flags which happen to "fix" the bug, because it you do so, the bug "reappears" again when someone else runs it on a different computer, or you modify the code elsewhere.
Another less typical but still relevant class of bugs, which gets a lot of attention, is relying on things that the C standard tells you "not to use" by calling them undefined behavior. What they do can randomly change between compilers, compiler versions, compiler settings such as optimization level. Use only if you really know what you are doing. Relatively easy to find and fix because, well, it's all in the standard. Get a good reviewer, or use an AI tool.
-
Well in case if something works totally correct with -O2, but breaks with -O3, or even worse, works OK with -O1, but breaks with -O2, whom fault is that?
Depends. For example, you can read in my original post in the Microchip forum that I was immediately warned about a possibly missing "volatile". Such problems can easily screw the code under strong optimization. But here, the code wasn't at fault, which Microchip acknowledged. They created a SFIX, which means that a fix could be provided sometime later.
The real question is how much do you need stronger optimization. In my case, that was really optional as my program, which works flawlessly with -O2, occupies about 50% of the program space and runs fast enough.
If you really want strong optimization, you either need to know the limits of your dev environment, or simply write assembly language.
The listing below (which I published in the Microchip forum), shows several problems in this simple switch block.
First, the cycles/instructions counts don't make sense. How can you make relevant optimizations if you can't even count correctly ?
Second, unlike what the XC8 manual says, the #pragma switch won't override XC8's strategy whatever the optimization level. Another bug IMO.
Third, how can -O3 leave line 2024 ? I haven't checked with -O2 or -O1, but this is real nonsense to me.
1995 ; Switch size 1, requested type "simple"
1996 ; Number of cases is 7, Range of values is 0 to 6
1997 ; switch strategies available:
1998 ; Name Instructions Cycles
1999 ; direct_byte 20 6 (fixed)
2000 ; simple_byte 22 12 (average)
2001 ; jumptable 260 6 (fixed)
2002 ; Chosen strategy is simple_byte
2003 0352 3A00 xorlw 0 ; case 0
2004 0353 1903 skipnz
2005 0354 2B2C goto l341
2006 0355 3A01 xorlw 1 ; case 1
2007 0356 1903 skipnz
2008 0357 2B30 goto l343
2009 0358 3A03 xorlw 3 ; case 2
2010 0359 1903 skipnz
2011 035A 2B33 goto l2320
2012 035B 3A01 xorlw 1 ; case 3
2013 035C 1903 skipnz
2014 035D 2B38 goto l2322
2015 035E 3A07 xorlw 7 ; case 4
2016 035F 1903 skipnz
2017 0360 2B3D goto l346
2018 0361 3A01 xorlw 1 ; case 5
2019 0362 1903 skipnz
2020 0363 2B48 goto l2326
2021 0364 3A03 xorlw 3 ; case 6
2022 0365 1903 skipnz
2023 0366 2B4D goto l350
2024 0367 2B68 goto l342 ; <--- Wow, nice -O3 feature !!!
2025 0368 l342:
2027 ; if (!PORTCbits.RC5)
2028 0368 0020 movlb 0 ; select bank0
2029 0369 1A8E btfsc 14,5 ;volatile
2030 036A 2C38 goto l2410
-
Third, how can -O3 leave line 2024 ?
This is the sort of nonsense that the Hi-TECH Omniscient-Code-Generation(OCG) technology is "supposed to" remove. That it obviously does not at -O3 suggest that OCG is not meeting its press releases.
I suspect that this behavior is a result of how the ISO C specification is usually implemented for a switch case statement. The effects of the optimizations appears to be limited for the case statements within the "{}" braces of the switch statement. An execution path transfer statement is required to direct execution to the next block. Where that block is located seems opaque to the optimizer.
-
PC it's a common thing to not go beyond -O2 as things may break.
I thought the usual reason for avoiding -O3 was that it significantly increases code size for only a slight improvement in speed? (Inlining to a silly extent, for example.)
Although for this bug, apparently it also happens with -Os, which is quite commonly used?
-
I thought the usual reason for avoiding -O3 was that it significantly increases code size for only a slight improvement in speed? (Inlining to a silly extent, for example.)
Yes, the purpose of -O3 is exactly that it's the "extreme" speed optimization which can blow up compilation speed (many don't consider this relevant anymore; it was actually a significant argument for the different optimization levels 20-30 years ago), and the binary size, and sometimes it does poor choices which may reduce the speed (e.g., pipelining and caching of the target machine not understood perfectly by optimizer; e.g. inlining too much can kill performance if fetching the program produces more cache misses!) - so it's a setting one normally uses "under supervision", so, you measure the results - did you get the speed benefit you wanted, and was the program size increase acceptable.
-O2 and -Os are the most obvious defaults on most targets and projects. -O1 instead of -O2 if compilation speed is important, but it rarely is nowadays.
The "trying to avoid compiler bugs by not using -O3" is real but a more subtle thing - compiler bugs exist at any optimization level because to err is human, that's all to it; highest optimization levels only increase the probability somewhat because the compiler is doing more advanced things, so more surface area for bugs, but that's a rather weak correlation in the end.
The "avoiding -O3 because my own code is buggy and I don't want to admit it is and -O0 seems to fix it" is the sad story which we unfortunately see on forums. This way of thinking should be left behind, with one caveat: there is a real gray area of commonly used and useful practices which are technically against the standard, but make sense and existing software uses them - in such cases, compiler designer should consider real-world use above the standard-correctness and not break programs whose intent is sensible, clear and justified de-facto usage of language. But these are rare exceptions, too, it is unfortunately too easy to use this as an argument to rationalize one's own procrastination / unwillingness to fix one's own code when that would be the best solution to everyone - just fix the bug and make the code standard compliant.
-
Yes, the purpose of -O3 is exactly that it's the "extreme" speed optimization which can blow up compilation speed
GCC does crazy things in -O3. For example I've seen it unrolling loops even though this actually made the code slower.
XC8 has nothing to do with GCC - their compiler is totally proprietary. The optimized -O3 version wasn't free and was rather expensive. Now they made it free, much more people started using it, so all the bad stuff that went undetected is now coming out.
-
This brings one question: if I wanted to remove line 2024, how can I "split" the build sequence to modify the asm source before it gets assembled ?
-
XC8 has nothing to do with GCC - their compiler is totally proprietary.
Being "proprietary" is actually not true. Being based off GCC isn't true as of late, but I think this was with the early versions of XC8, and I mean "XC8" as it is called.
The proprietary version of Microchip's 8-bit compiler was mcc8, which itself was based off HI-TECH C if I'm not mistaken, which they had licensed or bought altogether, I'm not sure.
But with the release of XC8, then it became an entirely different code base, which again I think was first based off GCC (heavily modified). It's now based of LLVM, as mentioned in their recent release notes, under the LLVM Release License.
-
XC8 has nothing to do with GCC - their compiler is totally proprietary.
Being "proprietary" is actually not true. Being based off GCC isn't true as of late, but I think this was with the early versions of XC8, and I mean "XC8" as it is called.
The proprietary version of Microchip's 8-bit compiler was mcc8, which itself was based off HI-TECH C if I'm not mistaken, which they had licensed or bought altogether, I'm not sure.
But with the release of XC8, then it became an entirely different code base, which again I think was first based off GCC (heavily modified). It's now based of LLVM, as mentioned in their recent release notes, under the LLVM Release License.
Microchip bought HITECH and renamed their compiler to XC8. I don't think it was based on GCC (otherwise it would have to be open source)/ They didn't even had a linker and used what they called "omniscient" code generation and compiled stack (where all stack allocations are silently converted to static variables) because small PICs don't have a stack. It didn't ever support C99.
Later they have changed front-end to clang, added C99 support, but kept their old code generation.
-
And expect to see both bugs and inefficiency in a proprietary compiler for a relatively small CPU player (Microchip) on an ancient and unimportant architecture (8-bit PICs). If O3 was a special paid level earlier, I would not be surprised if the user base size was tiny and not much information of how that really works leaked out from the few companies that used it.
-
Frankly speaking, XC8 is TWO different compilers: one for the AVR architecture (based on GCC), the other for PIC (originally based on the HITECH compiler).
-
Relatively easy to find and fix because, well, it's all in the standard.
Absolutely trivial, you only have to commit to memory and carefully enforce over two hundred exceptions some of which are so obscure or oblique that even the people writing the compilers can't agree on how to manage them, or it takes you a dozen re-readings and an hour of testing on Godbolt to figure out what the text is actually talking about.
More than half of which are completely unnecessary because the hardware that the special-case was created for hasn't existed for 30 or 40 years, but the compiler still gets to use it to break your code because it allows the compiler writers to show you how much cleverer they are than you.
-
And expect to see both bugs and inefficiency in a proprietary compiler for a relatively small CPU player (Microchip) on an ancient and unimportant architecture (8-bit PICs).
What are you talking about? Microchip is one of the biggest MCU companies in the world, and 8-bit PICs are a big part of their sales. Listed on NASDAQ. All the information is public. Grab the reports and read them.
-
XC8 for AVR is definitely based on GCC.
Looking back at older release notes, XC8 for PIC was based on HI-TECH compilers. There were 2 different compilers, one for the older series of 8-bit PIC and one for the PIC18 series, both taken from HI-TECH and rebadged. I seemed to remember a switch to GCC, but I don't know if that's something I had heard they planned to do (before deciding on LLVM), that's quite possible.
I don't know how many modifications they made to the original HI-TECH compilers, if any, they made before XC8 2.0, which switched to LLVM.
So currently both variants of XC8 (maybe we can still call that 3, as the backend for the PIC18 series is likely still separate from the backend for other series) are based on open-source compilers. XC8 for AVR is entirely open source as far as I've seen. for XC8 PIC, I don't know the LLVM license well enough to figure out if they managed to keep the backends closed source - in any case, the front end should be completely open. Since the LLVM seems to be more "liberal" than that of GCC, it's quite possible they finally decided on LLVM not so much for its merits (they already had experience with GCC with XC16 and XC32), but just for licensing reasons, to keep the PIC backends closed source. Would be "interesting" to know if that was their main rationale.
Also curious to see if they're gonna switch to LLVM for XC16 and XC32 too in the future.
-
XC8 for AVR is definitely based on GCC.
This is just the old Atmel compiler for AVR, which they renamed to XC8 too. So now there's two XC8 (XC8 for PICs and XC8 for AVR), but they have nothing in common except the name. OP is referring to 8-bit PIC compiler, not the AVR one.
They did a similar thing with ATSAM - they now have two XC32 compilers - XC32 for PIC32M (MIPS based) and XC32 for ATSAM (ARM32 based). They also renamed newer ATSAM chips (which they designed after Atmel acquisition) into PIC32C. Worse yet, they added the third XC32 compiler for PIC32A (which are the same as dsPIC33AK), but later they renamed this compiler into XC-DSC.
If you ask me, all these renamings don't do them any good.
-
What about -Os? This is the most common and most size-efficient mode.
-
If you ask me, all these renamings don't do them any good.
Yup, it's like the marketing dept wanted a nice clean toolset for each of the 8/16/32 bit mcu's.
But it's just superficial and ends up causing technical confusion, as seen above.
-
Quick update: Microchip told me that the problems discussed here will be fixed in the next XC8 release, which is currently being tested. No idea if this will fix the instructions/cycles counts in the switch strategies and whether the #pragma switch will allow overriding the strategy choosen by the compiler. Stay tuned !