Author Topic: Float vs fix math  (Read 7744 times)

0 Members and 1 Guest are viewing this topic.

Offline viperTopic starter

  • Frequent Contributor
  • **
  • Posts: 301
  • Country: us
Float vs fix math
« on: October 27, 2025, 09:42:03 pm »
Using an ESP32-S3, which has some ability with floats but literally everything I have learned says to find work arounds to avoid FP which possible, and FP has several potential hang ups like rounding and such. 

I'm not saying I am firm on one side or the other, and possible I will have no choice but to use some FP as I do some tougher DSP work, but sounds like a possible considerable time penalty for using it. 

There is nothing in my project that really needs both high precision and speed.  It would be one or the other. 
 

Offline Doctorandus_P

  • Super Contributor
  • ***
  • Posts: 5346
  • Country: nl
Re: Float vs fix math
« Reply #1 on: October 27, 2025, 10:17:05 pm »
... as I do some tougher DSP work, but sounds like a possible considerable time penalty for using it. 

In that case, you really should do some studying to get to know the differences between those two approaches. And it's not an easily answered topic. On an 8-bitter 32 bit integer calculations can be slower then floating point operations, there are lots of microcontrollers with built in floating point instructions. "real DSP's" have all sorts of extras to speed up specific types of calculations, etc.

For many "hobby level" projects, you can just choose a uC family, and then stick with it for all projects, but if you or your projects outgrow that, then there are a gazillion of choices these days. I've been happy for 15+ years myself with the (previous) Atmel AVR's, (Overall CPU usage <0.1%) until I wanted to experiment with TFT LCD's. Then I suddenly wanted an uC that could shove data around at least 10x faster then the old AVR's.

I suspect that you will get different and maybe even contradicting advise in this thread, and it will not be of much use without a lot of (self) study.
 
The following users thanked this post: viper, mskeete

Offline Psi

  • Super Contributor
  • ***
  • Posts: 12634
  • Country: nz
Re: Float vs fix math
« Reply #2 on: October 27, 2025, 10:25:04 pm »
Do a test, write a simple DSP example and see how fast you can run it. Then compare the speed of that to the expected complexity of your actual design to get a rough idea how fast it will run.

Using floats on an 8 bit mcu running at 20mhz works fine to do a fair number of float calculations, more than you think, mixed in with general integer stuff.
So a 32bit ESP32-S3 running at 240mhz will be orders of magnitude better than that.

I would try floats first.
« Last Edit: October 27, 2025, 10:33:17 pm by Psi »
Greek letter 'Psi' (not Pounds per Square Inch)
 
The following users thanked this post: viper

Offline Psi

  • Super Contributor
  • ***
  • Posts: 12634
  • Country: nz
Re: Float vs fix math
« Reply #3 on: October 27, 2025, 10:34:52 pm »
Actually, I just checked, the ESP32-S3 has a single precision floating point unit, it's not doing it in software, so speed wont be an issue.

Obviously it does depend how accurate you need your floating calcs to be.
« Last Edit: October 27, 2025, 10:41:39 pm by Psi »
Greek letter 'Psi' (not Pounds per Square Inch)
 

Online langwadt

  • Super Contributor
  • ***
  • Posts: 5783
  • Country: dk
Re: Float vs fix math
« Reply #4 on: October 27, 2025, 10:40:00 pm »
Do a test, write a simple DSP example and see how fast you can run it. Then compare the speed of that to the expected complexity of your actual design to get a rough idea how fast it will run.

Using floats on an 8 bit mcu running at 20mhz works fine to do a fair number of float calculations, more than you think, mixed in with general integer stuff.
So a 32bit ESP32-S3 running at 240mhz will be orders of magnitude better than that.

I would try floats first.

and since the ESP32-S3 afaict only has a single precision FPU, keep that in mind when programming in C, unless you need doubles make sure to specify constants as floats and use the float math functions, the default is doubles (I think the AVR GCC cheats and defaults to floats)


 

Offline Psi

  • Super Contributor
  • ***
  • Posts: 12634
  • Country: nz
Re: Float vs fix math
« Reply #5 on: October 27, 2025, 10:43:58 pm »
I'm not actually sure, if you do double precision float math on a MCU with only a single precision FPU, does it use the FPU to help with double precision calculations or does it fall back to pure int math and not use the FPU at all? One would hope it can use it to accelerate the double precision cals somehow
Greek letter 'Psi' (not Pounds per Square Inch)
 

Online hans

  • Super Contributor
  • ***
  • Posts: 1966
  • Country: 00
Re: Float vs fix math
« Reply #6 on: October 27, 2025, 10:46:22 pm »
Indeed, float vs integer can very much depend which operations are being carried out.
Even simple 32-bit integers don't need to be incredibly slow on small micros.. as long as the calculations stay inside the addition/subtraction domain. Unfortunately, often times some multiplication gets involved as thats how high dynamic range is quickly utilized.

Similarly, floats shine in the logarithmic domain, like calculating the reciprocal.. All you need to do is multiply the FP exponent by *-1, which can be done with a 2x 8-bit integer subtractions. In integer domain, you need to do long tail division or beg your MCU has at least an integer divider which is faster than half a dozen cycles.

Or square root. A very very rough approximation would be to simply divide by the exponent by 2 and don't even bother correcting the mantissa:
https://bits.stephan-brumme.com/squareRoot.html
And there is the famous fast inverse square root example..

Unfortunately, some people get principal about using floats on micro's. A while ago I got nearly harassed on Reddit /r/embedded for using floats on an AVR temperature sensor. It was running at 1 sample per second. There were no power constraints. The floating point library took half the codespace of the ATMEGA328P, which means it fits and runs. It may be a code smell for such application, but not all projects need to be able to go to the moon, and for a hobby project with minimal time investment it put down all the checkmarks.

I'm not actually sure, if you do double precision float math on a MCU with only a single precision FPU, does it use the FPU to help with double precision calculations or does it fall back to pure int math and not use the FPU at all? One would hope it can use it to accelerate the double precision cals somehow

It would be converting between `double` FP64 and `float` FP32 a lot.
This is a trap with floats. E.g. cos() and sin() are the double FP64 variants in <math.h>, while cosf and sinf are the float FP32 variants. Take a good look at which routines/instructions get used for which operations, as one or the other may incur a lot of useless FP32/FP64 conversions.
E.g. https://godbolt.org/z/sno11Eebc
« Last Edit: October 27, 2025, 10:51:53 pm by hans »
 

Offline Psi

  • Super Contributor
  • ***
  • Posts: 12634
  • Country: nz
Re: Float vs fix math
« Reply #7 on: October 27, 2025, 10:53:50 pm »
Unfortunately, some people get principal about using floats on micro's. A while ago I got nearly harassed on Reddit /r/embedded for using floats on an AVR temperature sensor. It was running at 1 sample per second. There were no power constraints. The floating point library took half the codespace of the ATMEGA328P, which means it fits and runs. It may be a code smell for such application, but not all projects need to be able to go to the moon, and for a hobby project with minimal time investment it put down all the checkmarks.

Agreed, I've been quite surprised with how many float calcs I can do in one of my ATMega64A products without causing any slowdown to the control loops or i/o loops.
(Most of the internal/external calc loops are either 50hz or 25hz).
The main loop has a float calc doing some sin/cos stuff, most of the IO has some float scaling/cal stuff going on per sample and there's a function doing exponential tracking of a current to PWM curve using floats.
Greek letter 'Psi' (not Pounds per Square Inch)
 

Offline nctnico

  • Super Contributor
  • ***
  • Posts: 30193
  • Country: nl
    • NCT Developments
Re: Float vs fix math
« Reply #8 on: October 27, 2025, 10:59:43 pm »
Using an ESP32-S3, which has some ability with floats but literally everything I have learned says to find work arounds to avoid FP which possible, and FP has several potential hang ups like rounding and such. 
Ignore the dogmas. Use floating point if you want unless you run out of processor power. It is simple like that. Being aware of rounding and resolution is necessary though. OTOH, code becomes so much easier to follow if variables represent SI units (like Volt, Ampere, meters, Hertz, rpm, Pascal, etc, etc).
« Last Edit: October 27, 2025, 11:04:56 pm by nctnico »
There are small lies, big lies and then there is what is on the screen of your oscilloscope.
 

Offline Psi

  • Super Contributor
  • ***
  • Posts: 12634
  • Country: nz
Re: Float vs fix math
« Reply #9 on: October 27, 2025, 11:03:24 pm »
Even when/if you do run out of processing power due to floats. It's going to be one specific block of float code that's using all the processing power and causing all the slowdown issues.
So best to just convert that block to integer math if needed. Not the whole thing.
« Last Edit: October 27, 2025, 11:20:20 pm by Psi »
Greek letter 'Psi' (not Pounds per Square Inch)
 
The following users thanked this post: nctnico

Offline viperTopic starter

  • Frequent Contributor
  • **
  • Posts: 301
  • Country: us
Re: Float vs fix math
« Reply #10 on: October 27, 2025, 11:07:58 pm »
The most intense part here is the DSP filter that will run a calculation.  It could be around 1M multiplies to do every 5 sec or so.  I am also realizing the heart of these equations for the DSP rely on floating math, so I may have little choice, or at least I could prob do some conversions. 

I appreciate and agree this is hard to really 'answer' without a specific use case.  Was just trying to get ideas in my head as I work if I should be thinking of ways to use integers for this and that. 
 

Offline nctnico

  • Super Contributor
  • ***
  • Posts: 30193
  • Country: nl
    • NCT Developments
Re: Float vs fix math
« Reply #11 on: October 27, 2025, 11:20:11 pm »
Many years ago I implemented a DSP algorithm which started life as code which uses floating point. I optimised it using fixed point until it was fast enough. Not more.
There are small lies, big lies and then there is what is on the screen of your oscilloscope.
 

Offline ejeffrey

  • Super Contributor
  • ***
  • Posts: 4841
  • Country: us
Re: Float vs fix math
« Reply #12 on: October 27, 2025, 11:44:33 pm »
Using an ESP32-S3, which has some ability with floats but literally everything I have learned says to find work arounds to avoid FP which possible, and FP has several potential hang ups like rounding and such. 

I'm not saying I am firm on one side or the other, and possible I will have no choice but to use some FP as I do some tougher DSP work, but sounds like a possible considerable time penalty for using it. 

There is nothing in my project that really needs both high precision and speed.  It would be one or the other.

As has been widely explained above, if you have a processor with a floating point unit, floating point math is going to be about as fast as fixed point if not faster.

So for general applications it's usually easier to just use floating point if you have an FPU available.

There are specific DSP algorithms that only work well with fixed point, and others that have desirable properties when used with fixed point.  For instance, some algorithms rely on exact cancellation of rounding errors, or specific overflow behavior.  Because the truncation of floating point is data dependent, your filter parameters can vary slightly with your input data.  In a lot of cases this variation is totally inconsequential, and the better average dynamic range is only a benefit.  In some cases it matters quite a bit.

Mostly if you need to use fixed point for one of these reasons, you will know it.  If you are just learning and not trying to absolutely maximize your dynamic range, don't get too hung up on this.  Just use floating point unless you have a specific reason not to.  Even if you are going to want fixed point eventually usually it makes sense to first prototype with floating point until you are happy with the general operation.  Then do the optimized design which might involve transitioning to fixed point.
 

Online radiolistener

  • Super Contributor
  • ***
  • Posts: 5741
  • Country: Earth
Re: Float vs fix math
« Reply #13 on: October 28, 2025, 12:45:00 am »
The choice of number representation really depends on your hardware and project requirements. Floating point has lower precision and resolution for small signals at the same bit width, but it handles large dynamic ranges easily. Fixed-point is more efficient and faster on hardware without an FPU, but you may need more bits to cover the same range.

On FPGA, using floating-point is generally challenging due to implementation complexity, whereas on modern CPU/MCU, floating-point is typically supported in hardware. Therefore, the choice largely depends on the available resources.

If you convert a filter with float coefficients to fixed-point, the filter parameters will change slightly due to rounding, so this should be considered during the filter design stage.
 

Offline Geoff-AU

  • Frequent Contributor
  • **
  • Posts: 397
  • Country: au
Re: Float vs fix math
« Reply #14 on: October 28, 2025, 01:28:59 am »
literally everything I have learned says to find work arounds to avoid FP which possible

As others have said, that "rule" is obsolete dogma these days.  It used to be necessary with 8-bit microcontrollers running around 1MHz that had to do the calculation in software.  Larger/modern microcontrollers (like ESP32) have floating point units (FPU) to offload the work to, and clock rates around 80MHz.

I remember having to work something out in real time and just for giggles I broke the rule and threw a floating point calc into ISR.  It was an ESP32 and my ISR only used up about 5% of the available execution time (toggle a GPIO to profile the time spent).  We're spoilt these days!
 

Offline iMo

  • Super Contributor
  • ***
  • Posts: 6898
  • Country: li
Re: Float vs fix math
« Reply #15 on: October 28, 2025, 08:21:16 am »
Afaik the naked fp single precision multiply on the ESP32 takes 4cycles, with some overhead say 8cycles.
At 80MHz clock that is say 100ns. One million mults is then around 0.1sec..
Readers discretion is advised..
 

Online Jeroen3

  • Super Contributor
  • ***
  • Posts: 4575
  • Country: nl
  • Embedded Engineer
    • jeroen3.nl
Re: Float vs fix math
« Reply #16 on: October 28, 2025, 08:32:53 am »
Afaik the naked fp single precision multiply on the ESP32 takes 4cycles, with some overhead say 8cycles.
At 80MHz clock that is say 100ns. One million mults is then around 0.1sec..
Though technically correct, you should consider that you need to copy the data to and back out from the FPU.
So it will be at minimum 2 copies, 4 cycles, 1 copy. Especially if you can create cache-misses you can get significant slowdowns that are difficult to track down.

In the end using fixed point the time is lost doing the checks around and the additional scalings.
floating point is a bit easier, but less deterministic.

You may still need to run a few isfinites().

I'd recommend to create your algorithm and do benchmark on-target.
Fixed point will require more planning, but it could still be more performant or deterministic.

If you are constantly changing the alrogrithm and there is "plenty" of time, I'd use the floating point. It's just easier to work with.
 

Offline AVI-crak

  • Regular Contributor
  • *
  • Posts: 151
  • Country: ru
    • Rtos
Re: Float vs fix math
« Reply #17 on: October 28, 2025, 09:49:48 am »
#define PI      3.14159274101f   /// 0x40490fdb
#define Pi      3.14159250259f   /// 0x40490fda
#define Pi_d    3.1415926535897932384626433832795
 

Offline Peabody

  • Super Contributor
  • ***
  • Posts: 2770
  • Country: us
Re: Float vs fix math
« Reply #18 on: October 28, 2025, 02:42:35 pm »
I'm new to fixed point, and currently working on a home telephone project that requires DTMF detection using the Goertzel algorithm.  I'd like to get it to work with an Atmega328P (a Nano), and of course AVR is just 8-bit with no FPU.  However, I think it does have signed and unsigned 8-bit multiply instructions.  The limiting factor on this project is code size, not ram.  So I will be trying both a floating point version and a fixed point version to see how much difference there is.

The Goertzel algorithm is essentially a single-frequency FFT calculation.  It uses cos(), division, sqrt(), and lots of multiplication.  However, the cos() and division can be avoided by pre-computing the eight DTMF frequency coefficients and hard-coding them.  And sqrt() is only required to determine the magnitude of the signal to compare to the pre-set threshold, and magnitude squared would work just as well, compared to a squared threshold.

So I'm not sure there will be any significant difference.  Speed isn't really important in this case.  I guess multiplying two floats requires more code that multiplying two ints, but they are going to be functions in either case, plus the fixed point version will also need a >>14 loop.  So it may come down to just a few bytes difference.  Based on what I've done so far, floats are easier to deal with. With fixed point you have to worry about overruns.
 

Offline nctnico

  • Super Contributor
  • ***
  • Posts: 30193
  • Country: nl
    • NCT Developments
Re: Float vs fix math
« Reply #19 on: October 28, 2025, 02:46:43 pm »
I'm new to fixed point, and currently working on a home telephone project that requires DTMF detection using the Goertzel algorithm.  I'd like to get it to work with an Atmega328P (a Nano), and of course AVR is just 8-bit with no FPU.  However, I think it does have signed and unsigned 8-bit multiply instructions.  The limiting factor on this project is code size, not ram.  So I will be trying both a floating point version and a fixed point version to see how much difference there is.
Side node: Goertzel is useless for DTMF decoding. A better approach is to use two band filters and use a frequency counter (zero crossing detection) to measure the frequencies. This will give you a much more accurate and better defined DTMF decoder. Make sure the sampling rate is >=16kHz (either sampled directly or use oversampling before filtering).
There are small lies, big lies and then there is what is on the screen of your oscilloscope.
 

Offline Peabody

  • Super Contributor
  • ***
  • Posts: 2770
  • Country: us
Re: Float vs fix math
« Reply #20 on: October 28, 2025, 03:08:33 pm »
Well so far Goertzel is working fine for DTMF.  It appears to be what most people use for DTMF in the Arduino world.
 

Offline ejeffrey

  • Super Contributor
  • ***
  • Posts: 4841
  • Country: us
Re: Float vs fix math
« Reply #21 on: October 28, 2025, 05:33:10 pm »
It's widely used for DTMF decoding across the board and is proven to work well.  Claiming it's "useless" is just nonsense.
 

Offline nctnico

  • Super Contributor
  • ***
  • Posts: 30193
  • Country: nl
    • NCT Developments
Re: Float vs fix math
« Reply #22 on: October 28, 2025, 06:56:02 pm »
Well so far Goertzel is working fine for DTMF.  It appears to be what most people use for DTMF in the Arduino world.
That could be but early in my career one of my mentors (a very smart and experienced engineer) pointed out the flaws in using Goertzel for DTMF decoding (sensitivity and selectivity which means you can't meet the DTMF specs) so we choose to go for the frequency counter route. The fact that many people use a method doesn't mean it is the right one. For the project I worked on at that time, the DTMF decoding needed to work very reliably in commercial call center applications.
« Last Edit: October 28, 2025, 08:23:31 pm by nctnico »
There are small lies, big lies and then there is what is on the screen of your oscilloscope.
 

Offline viperTopic starter

  • Frequent Contributor
  • **
  • Posts: 301
  • Country: us
Re: Float vs fix math
« Reply #23 on: October 28, 2025, 06:58:48 pm »
I have no test data other than specs for this ESP unit which shows 238MOPS for FPU multiply?  240mhz.  That is 4 cycles, single precision, but I hear you guys on considering other operations.  I'm just trying to get an idea if this is a few mS or a few minutes.   :scared:

There is also SIMD, which I have absolutely no experience with.  Basically get the engine running decent, then tune to the moon later. 
« Last Edit: October 28, 2025, 07:00:55 pm by viper »
 

Online peter-h

  • Super Contributor
  • ***
  • Posts: 6010
  • Country: gb
  • Doing electronics since the 1960s...
Re: Float vs fix math
« Reply #24 on: October 28, 2025, 09:19:00 pm »
Re GCC and single floats, I don't know about "all GCC" but for sure arm32 GCC uses doubles, which are then (IME trig measurements) some 20x slower than single floats, on a single float hardware CPU.
Z80 Z180 Z280 Z8 S8 8031 8051 H8/300 H8/500 80x86 90S1200 32F417
 


Share me

Digg  Facebook  SlashDot  Delicious  Technorati  Twitter  Google  Yahoo
Smf

 

-->