The rosetta version *did* run 10% faster than the native version (wall clock time), even including translation overhead, and got the same printed results. There's no arguing against that.
This is all far beyond my knowledge, but I did read of an Apple Silicon feature that may relate to this result.
Intel processors have a strong memory ordering model that adds implicit memory barriers. ARM processors have a weaker memory ordering model to allow more reordering for better performance. This would make Rosetta performance terrible if it needed to emulate this behaviour with additional barrier instructions to accurately emulate an Intel processor.
To solve this, Apple built a TSO (Total Store Ordering) mode that changes the processor’s memory ordering model to be similar to Intel’s. This allows Rosetta to translate x86-64 load and store instructions to ARM directly.
This is a fairly obvious idea -- obvious enough that RISC-V started the process to add it as an option about 3 1/2 years ago and ratified it as a standard in July 2018.
You can build compliant RISC-V cores that use RVWMO (RISC-V Weak Memory Ordering, which is actually somewhat stronger than ARM's, but still allows high performance) or RVTSO (RISC-V Total Store Ordering), or that are switchable between the two.
Software written for RVWMO will run correctly on any CPU or mode, but may run somewhat slower is TSO is enabled. Software written for RVTSO must run on a CPU implementing RVTSO full-time on in RVTSO mode on a CPU that is switchable.
RVWMO was designed by a panel of world experts on memory consistency after bugs were found in the original RISC-V memory consistency specification in 2016.
Here's a discussion presentation from mid 2017, with a lot of background:
https://www.bsc.es/sites/default/files/public/u1810/arvind_0.pdfAnd a status update at the end of 2017
https://riscv.org/wp-content/uploads/2017/12/Tue0954-RISC-V_Memory_Model-Lustig.pdfIt seems that TSO mode may only be available on the high performance Firestorm cores. Perhaps the ARM version started on the low power cores and was moved to the high performance cores. Is there a way to force the program to run on the high performance cores?
If the Rosetta version can only run on the high performance cores, this may explain the speed discrepancy.
I have no idea about a difference in the cores, but any program that runs CPU-bound for more than maybe 100mS (or even 1mS?) causes a transition to a high performance core.