Indeed, float vs integer can very much depend which operations are being carried out.
Even simple 32-bit integers don't need to be incredibly slow on small micros.. as long as the calculations stay inside the addition/subtraction domain. Unfortunately, often times some multiplication gets involved as thats how high dynamic range is quickly utilized.
Similarly, floats shine in the logarithmic domain, like calculating the reciprocal.. All you need to do is multiply the FP exponent by *-1, which can be done with a 2x 8-bit integer subtractions. In integer domain, you need to do long tail division or beg your MCU has at least an integer divider which is faster than half a dozen cycles.
Or square root. A very very rough approximation would be to simply divide by the exponent by 2 and don't even bother correcting the mantissa:
https://bits.stephan-brumme.com/squareRoot.htmlAnd there is the famous fast inverse square root example..
Unfortunately, some people get principal about using floats on micro's. A while ago I got nearly harassed on Reddit /r/embedded for using floats on an AVR temperature sensor. It was running at 1 sample per second. There were no power constraints. The floating point library took half the codespace of the ATMEGA328P, which means it fits and runs. It may be a code smell for such application, but not all projects need to be able to go to the moon, and for a hobby project with minimal time investment it put down all the checkmarks.
I'm not actually sure, if you do double precision float math on a MCU with only a single precision FPU, does it use the FPU to help with double precision calculations or does it fall back to pure int math and not use the FPU at all? One would hope it can use it to accelerate the double precision cals somehow
It would be converting between `double` FP64 and `float` FP32 a lot.
This is a trap with floats. E.g. cos() and sin() are the double FP64 variants in <math.h>, while cosf and sinf are the float FP32 variants. Take a good look at which routines/instructions get used for which operations, as one or the other may incur a lot of useless FP32/FP64 conversions.
E.g.
https://godbolt.org/z/sno11Eebc