Author Topic: Apples new M1 microprocessor  (Read 88552 times)

0 Members and 11 Guests are viewing this topic.

Offline brucehoult

  • Super Contributor
  • ***
  • Posts: 6458
  • Country: nz
Re: Apples new M1 microprocessor
« Reply #75 on: November 16, 2020, 03:08:24 pm »
Precisely. It's a 1970s design which has been hacked and hacked and hacked.

ARM however is a 1980s design which has been hacked and hacked and hacked.

Aarch64, which is what Apple's chips run exclusively, is nearly totally different to the 1980s 32 bit ARM design. It is very much a clean sheet post-2000 design, having learned things (good and bad) from PowerPC and Alpha and Itanium. It is much closer to Alpha and to RISC-V than to 1980s ARM.
 

Offline bd139

  • Super Contributor
  • ***
  • Posts: 23102
  • Country: gb
Re: Apples new M1 microprocessor
« Reply #76 on: November 16, 2020, 03:09:11 pm »
so then the question becomes: what works better - a native x86 core or an emulation running on a different architecture?   My money would be on the native core every time, all else being equal.

Turns out transpiling it from x86 to ARM actually might end up faster  :popcorn:

https://wccftech.com/m1-macbook-air-x86-emulation-faster-all-mac-models/

Note that includes i9 and Xeon CPUs that are shipping.

In that case, all else is not equal!   An emulation running on a brilliant processor will beat a slow native processor...    but if we are talking about two equally brilliant processors...  ?    My money is on native!

It’s not an emulation. It is directly recompiling the x86 code as ARM at startup. It treats x86 as a compiler IL.

Precisely. It's a 1970s design which has been hacked and hacked and hacked.

ARM however is a 1980s design which has been hacked and hacked and hacked.

Aarch64, which is what Apple's chips run exclusively, is nearly totally different to the 1980s 32 bit ARM design. It is very much a clean sheet post-2000 design, having learned things (good and bad) from PowerPC and Alpha and Itanium. It is much closer to Alpha and to RISC-V than to 1980s ARM.


I was mostly referring to RISC CPUs in general there.
« Last Edit: November 16, 2020, 03:10:58 pm by bd139 »
 

Offline Cerebus

  • Super Contributor
  • ***
  • Posts: 10576
  • Country: gb
Re: Apples new M1 microprocessor
« Reply #77 on: November 16, 2020, 03:10:41 pm »
so then the question becomes: what works better - a native x86 core or an emulation running on a different architecture?   My money would be on the native core every time, all else being equal.

Turns out transpiling it from x86 to ARM actually might end up faster  :popcorn:

https://wccftech.com/m1-macbook-air-x86-emulation-faster-all-mac-models/

Note that includes i9 and Xeon CPUs that are shipping.

In that case, all else is not equal!   An emulation running on a brilliant processor will beat a slow native processor...    but if we are talking about two equally brilliant processors...  ?    My money is on native!

transpiling \$\neq\$ emulating
Anybody got a syringe I can use to squeeze the magic smoke back into this?
 

Offline SilverSolder

  • Super Contributor
  • ***
  • Posts: 6126
  • Country: 00
Re: Apples new M1 microprocessor
« Reply #78 on: November 16, 2020, 03:40:46 pm »
It’s not an emulation. It is directly recompiling the x86 code as ARM at startup. It treats x86 as a compiler IL.

Ah, I missed that, so the "hit" is taken on startup?  - or is it perhaps only done once, and the resulting binary stored for later use?

In order to be faster than an x86 hardware implementation,  the target processor has to be faster at running its transpiled object code than the x86 processor is at running its native code... 

That just doesn't sound right, unless the hardware x86 implementation is really handicapped and unfixable... and nothing is unfixable, given clever enough engineers!  :D


 

Offline bd139

  • Super Contributor
  • ***
  • Posts: 23102
  • Country: gb
Re: Apples new M1 microprocessor
« Reply #79 on: November 16, 2020, 03:54:51 pm »
It's done once at first startup.

Converting x86 to ARM is interesting. x86 instructions are actually heavily microcoded so it's possible if you understand the semantics further than just the instructions but what side effects are, to produce more efficient code than the original architecture. On top of that some of the ARM instructions are super powerful. Nice example is CSEL, so to do a "max" function, you can:

CMP w0, w1                   # compare w0, w1
CSEL  w0, w0, w1, ge     # if w0 > w1 then store w0 in w0 else w1

That takes more instructions and more cycles in the x86 version. If you're calling that stuff thousands of times then you're on to a winner. But if you can reason about the compare and select function using an optimiser (peephole optimisers used to be used for this but I'm sure they have better ways) then you can shave off a feck load of microinstructions.

Fine example of this is the objective C dispatch stuff in macOS which is easy to reason about and heavily optimised per architecture so probably has a huge gain.
 

Online Berni

  • Super Contributor
  • ***
  • Posts: 5381
  • Country: si
Re: Apples new M1 microprocessor
« Reply #80 on: November 16, 2020, 03:58:08 pm »
It’s not an emulation. It is directly recompiling the x86 code as ARM at startup. It treats x86 as a compiler IL.

Ah, I missed that, so the "hit" is taken on startup?  - or is it perhaps only done once, and the resulting binary stored for later use?

In order to be faster than an x86 hardware implementation,  the target processor has to be faster at running its transpiled object code than the x86 processor is at running its native code... 

That just doesn't sound right, unless the hardware x86 implementation is really handicapped and unfixable... and nothing is unfixable, given clever enough engineers!  :D

Its proabobly the same as JIT compiling that has gotten so popular these days in the form of .Net and JavaScript. It crunches a block of code into native machine language and places a jump back into the compiler at the end of it, once the block of code gets to the end the JIT compiler takes over and compiles more of it. So code only gets compiled once its first executed. When it executes again its already native machine code, if it never executes then little/no work is wasted compiling it because it never gets there.

This was already used back in 2000 in the form of so called "High level emulation" where a Nintendo 64 emulator called UltraHLE used this aproach to get a massive performance boost in emulating the console on x86. Back then PCs ware generaly not fast enugh to emulate consoles this powerful, especialy something like the N64 with its very different 64bit instruction set, weird float math, proprietary GPU etc..

The issue with this on the fly translation aproach is that it can be very sensitive to the kind of code you want it to run. Some things translate really well and you get unbelevably good performance, while some code it just falls on its face and ends up running so slow to border on being unusable. If the code pulls some funky tricks the translator might really get confused about what it is trying to do and cause the resulting translated code to crash when run. So the results can varry wldly from app to app.
 

Online tszaboo

  • Super Contributor
  • ***
  • Posts: 9821
  • Country: nl
  • Current job: ATEX product design
Re: Apples new M1 microprocessor
« Reply #81 on: November 16, 2020, 04:11:10 pm »
so then the question becomes: what works better - a native x86 core or an emulation running on a different architecture?   My money would be on the native core every time, all else being equal.

Turns out transpiling it from x86 to ARM actually might end up faster  :popcorn:

https://wccftech.com/m1-macbook-air-x86-emulation-faster-all-mac-models/

Note that includes i9 and Xeon CPUs that are shipping.
In geekbench. It has been proved several times, that apple get points on geekbench that are completely bogus and manipulated. For example the Samsung whatever became faster than the iphone in geekbench x version, than x+1 came out and apple was again faster.
But you dont have to believe me, there is this guy's opinion, ha has an operating system named after him(although this was a few years ago.):
https://www.realworldtech.com/forum/?threadid=136526&curpostid=136666
 

Offline bd139

  • Super Contributor
  • ***
  • Posts: 23102
  • Country: gb
Re: Apples new M1 microprocessor
« Reply #82 on: November 16, 2020, 04:14:44 pm »
Geekbench is a load of shit for sure. The key thing though is measured compute performance. It appears that the answers are now in with Cinebench coming in...



I'm going to fluffle my shitty little Ryzen T495s laptop (3500U) which turns up 876 on single thread on R23, 3544 on multiple core and cost 470 quid new  :-DD. I'm not even going to try it on my big boy desktop  :popcorn:
 

Offline magic

  • Super Contributor
  • ***
  • Posts: 8061
  • Country: pl
Re: Apples new M1 microprocessor
« Reply #83 on: November 16, 2020, 04:23:17 pm »
CMP w0, w1                   # compare w0, w1
CSEL  w0, w0, w1, ge     # if w0 > w1 then store w0 in w0 else w1

That takes more instructions and more cycles in the x86 version.
You need to catch up with Pentium 4 ISA extensions ;)
Code: [Select]
00000000 <max>:
   0:   8b 54 24 04             mov    0x4(%esp),%edx
   4:   8b 44 24 08             mov    0x8(%esp),%eax
   8:   39 d0                   cmp    %edx,%eax
   a:   0f 4c c2                cmovl  %edx,%eax
   d:   c3                      ret   

BTW, I'm not sure if the jmp/mov version isn't faster when the outcome of comparison is repetitive and can be evaded by branch prediction.
 

Offline SilverSolder

  • Super Contributor
  • ***
  • Posts: 6126
  • Country: 00
Re: Apples new M1 microprocessor
« Reply #84 on: November 16, 2020, 04:24:31 pm »
It's done once at first startup.

Converting x86 to ARM is interesting. x86 instructions are actually heavily microcoded so it's possible if you understand the semantics further than just the instructions but what side effects are, to produce more efficient code than the original architecture. On top of that some of the ARM instructions are super powerful. Nice example is CSEL, so to do a "max" function, you can:

CMP w0, w1                   # compare w0, w1
CSEL  w0, w0, w1, ge     # if w0 > w1 then store w0 in w0 else w1

That takes more instructions and more cycles in the x86 version. If you're calling that stuff thousands of times then you're on to a winner. But if you can reason about the compare and select function using an optimiser (peephole optimisers used to be used for this but I'm sure they have better ways) then you can shave off a feck load of microinstructions.

Fine example of this is the objective C dispatch stuff in macOS which is easy to reason about and heavily optimised per architecture so probably has a huge gain.

The MAX example is interesting.  The transpiler would have to recognize the string of multiple x86 instructions in the source code that are implementing a "MAX" function, and reduce those instructions to a single instruction on the target -   Sounds difficult and prone to problems, since various source files could implement this in many different ways?

But what's to stop an x86 engineer from using similar tactics in a native x86? - including adding an extra MAX instruction, if that makes a big difference in real world applications?

Basically what I'm struggling with is the idea that one Turing machine architecture is intrinsically faster than another.  Basically, for this to be true, one of the two machines would have to be doing unnecessary work, essentially wasting time instead of executing the programmed instructions?  (given that they are equally fast!)
« Last Edit: November 16, 2020, 04:29:03 pm by SilverSolder »
 

Offline bd139

  • Super Contributor
  • ***
  • Posts: 23102
  • Country: gb
Re: Apples new M1 microprocessor
« Reply #85 on: November 16, 2020, 04:25:52 pm »
CMP w0, w1                   # compare w0, w1
CSEL  w0, w0, w1, ge     # if w0 > w1 then store w0 in w0 else w1

That takes more instructions and more cycles in the x86 version.
You need to catch up with Pentium 4 ISA extensions ;)
Code: [Select]
00000000 <max>:
   0:   8b 54 24 04             mov    0x4(%esp),%edx
   4:   8b 44 24 08             mov    0x8(%esp),%eax
   8:   39 d0                   cmp    %edx,%eax
   a:   0f 4c c2                cmovl  %edx,%eax
   d:   c3                      ret   

BTW, I'm not sure if the jmp/mov version isn't faster when the outcome of comparison is repetitive and can be evaded by branch prediction.

I get this with clang and optimiser doing its thing...

Code: [Select]
        mov     eax, esi
        cmp     edi, esi
        cmovge  eax, edi
 

Offline borjam

  • Supporter
  • ****
  • Posts: 911
  • Country: es
  • EA2EKH
Re: Apples new M1 microprocessor
« Reply #86 on: November 16, 2020, 04:29:41 pm »
Basically what I'm struggling with is the idea that one Turing machine architecture is intrinsically faster than another.  Basically, for this to be true, one of the two machines would have to be doing unnecessary work, essentially wasting time instead of executing the programmed instructions?
Nowadays it's harder to fathom because all people are seeing are the same CPUs from a single family: Intel.

In the past, when there was cutthroat competition between the Alpha, MIPS, SPARC, PA/RISC and POWER architectures you saw the crown jumping from one to the next every few months. And yes, I haven't mentioned Intel. They were producing mostly overpriced toys.

Have CPUs become perfect so that they can't improve anymore? I don't think so. Moreover, efficiency is really important as well.
 

Offline bd139

  • Super Contributor
  • ***
  • Posts: 23102
  • Country: gb
Re: Apples new M1 microprocessor
« Reply #87 on: November 16, 2020, 04:33:11 pm »
It's done once at first startup.

Converting x86 to ARM is interesting. x86 instructions are actually heavily microcoded so it's possible if you understand the semantics further than just the instructions but what side effects are, to produce more efficient code than the original architecture. On top of that some of the ARM instructions are super powerful. Nice example is CSEL, so to do a "max" function, you can:

CMP w0, w1                   # compare w0, w1
CSEL  w0, w0, w1, ge     # if w0 > w1 then store w0 in w0 else w1

That takes more instructions and more cycles in the x86 version. If you're calling that stuff thousands of times then you're on to a winner. But if you can reason about the compare and select function using an optimiser (peephole optimisers used to be used for this but I'm sure they have better ways) then you can shave off a feck load of microinstructions.

Fine example of this is the objective C dispatch stuff in macOS which is easy to reason about and heavily optimised per architecture so probably has a huge gain.

The MAX example is interesting.  The transpiler would have to recognize the string of multiple x86 instructions in the source code that are implementing a "MAX" function, and reduce those instructions to a single instruction on the target -   Sounds difficult and prone to problems, since various source files could implement this in many different ways?

But what's to stop an x86 engineer from using similar tactics in a native x86? - including adding an extra MAX instruction, if that makes a big difference in real world applications?

Basically what I'm struggling with is the idea that one Turing machine architecture is intrinsically faster than another.  Basically, for this to be true, one of the two machines would have to be doing unnecessary work, essentially wasting time instead of executing the programmed instructions?

Yes they could which is the point of having generic instruction sets and multi-pass optimisers and the like. They compile to something fairly generic (intermediate language - LLVM IR) and then specialist translators and optimisers translate that to ARM code and optimise it. I suspect Rosetta 2 just turns x86 back into LLVM IR and then recompiles it using the existing optimisers and translators. Well that's how I'd do it :). Apple contribute a lot to LLVM so it makes sense to leverage that. The key advantage is all your optimisation decisions are applicable to all code that is to be run whether it be from clang, rosetta or swift etc. If you come up with a speed hack one day or manage to compress a certain idiom into one instruction (common with ARM) then you can write an optimiser and everyone benefits instantly.

Regarding instructions, that's exactly why there are bazillions of instructions in x86-64 and the semantics around edge cases on some weird ones is how we get nice security holes.

Regarding different architectural speeds, turing completeness does not necessarily define something being efficient. In fact the original Universal Turing machine is horribly inefficent. Doesn't mean it's not useful or fast.

 

Offline bd139

  • Super Contributor
  • ***
  • Posts: 23102
  • Country: gb
Re: Apples new M1 microprocessor
« Reply #88 on: November 16, 2020, 04:53:46 pm »
Basically what I'm struggling with is the idea that one Turing machine architecture is intrinsically faster than another.  Basically, for this to be true, one of the two machines would have to be doing unnecessary work, essentially wasting time instead of executing the programmed instructions?
Nowadays it's harder to fathom because all people are seeing are the same CPUs from a single family: Intel.

In the past, when there was cutthroat competition between the Alpha, MIPS, SPARC, PA/RISC and POWER architectures you saw the crown jumping from one to the next every few months. And yes, I haven't mentioned Intel. They were producing mostly overpriced toys.

Have CPUs become perfect so that they can't improve anymore? I don't think so. Moreover, efficiency is really important as well.

That's where AMD are winning now.

As for Intel vs others, they mostly died because of hardware capital cost. Intel were far far far cheaper, the performance was largely better. I remember the company I worked for paying over £10k for one 360MHz PA-RISC CPU for an HP N-class whereas a 600MHz Pentium III Katmai was about the same speed for only £550. Buy two of them, a pile of ECC RAM, disks and windows NT and you had change left over from just the HP CPU. And only two years later they had 1.4GHz Tualatin parts that were 5x as fast (!).
 

Offline Cerebus

  • Super Contributor
  • ***
  • Posts: 10576
  • Country: gb
Re: Apples new M1 microprocessor
« Reply #89 on: November 16, 2020, 04:59:57 pm »

The MAX example is interesting.  The transpiler would have to recognize the string of multiple x86 instructions in the source code that are implementing a "MAX" function, and reduce those instructions to a single instruction on the target -   Sounds difficult and prone to problems, since various source files could implement this in many different ways?


It's easier than you might think. You scan the input instructions and build a tree structure based on their intent, you treat this exactly like you would treat the intermediate output of a compiler. At this point you have all the tricks of optimization and code generation in your hands. You can do dataflow analysis, you can reallocate registers and so on. Once you've done that you match the target instruction set to the intents signalled by the intermediate tree representation using pattern matching (or as Aho, Ganapathi and Tjang called it "Code Generation Using Tree Matching and Dynamic Programming"*). There's a good chance that the intent "find the larger of two values and store it" will be represented the same in the tree however the original source program represented it, if you have a pattern for a target instruction that matches that it will get generated.

You need to read the Aho et al paper and have a passing grasp of what compiler intermediate forms look like to really get a good feel for the method but it is both elegant and simple.

*ACM Transactions on Programming Languages and Systems, Vol. 11, No. 4, October 1989, Pages 491-516.
Anybody got a syringe I can use to squeeze the magic smoke back into this?
 
The following users thanked this post: bd139

Offline magic

  • Super Contributor
  • ***
  • Posts: 8061
  • Country: pl
Re: Apples new M1 microprocessor
« Reply #90 on: November 16, 2020, 05:20:24 pm »
I get this with clang and optimiser doing its thing...

Code: [Select]
        mov     eax, esi
        cmp     edi, esi
        cmovge  eax, edi
Okay, that's where ARM wins.
You no longer compare two words and maybe overwrite one of them, but now you pick one of two words and move it to a third register. This can't be done with CMOV because it's a two argument instruction like most of x86.
The reason those registers are allocated as they are is because of calling convention. The problem would go away if the function is inlined, assuming that code can be generated in such way that one of the inputs enters through the output register.

Good example how minor differences in implementations of almost the same instruction can make or break real world code.
 
The following users thanked this post: bd139

Offline SilverSolder

  • Super Contributor
  • ***
  • Posts: 6126
  • Country: 00
Re: Apples new M1 microprocessor
« Reply #91 on: November 16, 2020, 05:28:20 pm »

The MAX example is interesting.  The transpiler would have to recognize the string of multiple x86 instructions in the source code that are implementing a "MAX" function, and reduce those instructions to a single instruction on the target -   Sounds difficult and prone to problems, since various source files could implement this in many different ways?


It's easier than you might think. You scan the input instructions and build a tree structure based on their intent, you treat this exactly like you would treat the intermediate output of a compiler. At this point you have all the tricks of optimization and code generation in your hands. You can do dataflow analysis, you can reallocate registers and so on. Once you've done that you match the target instruction set to the intents signalled by the intermediate tree representation using pattern matching (or as Aho, Ganapathi and Tjang called it "Code Generation Using Tree Matching and Dynamic Programming"*). There's a good chance that the intent "find the larger of two values and store it" will be represented the same in the tree however the original source program represented it, if you have a pattern for a target instruction that matches that it will get generated.

You need to read the Aho et al paper and have a passing grasp of what compiler intermediate forms look like to really get a good feel for the method but it is both elegant and simple.

*ACM Transactions on Programming Languages and Systems, Vol. 11, No. 4, October 1989, Pages 491-516.

I think I get it - the "footprints in the snow" look the same no matter which order the steps were taken?
 

Offline SilverSolder

  • Super Contributor
  • ***
  • Posts: 6126
  • Country: 00
Re: Apples new M1 microprocessor
« Reply #92 on: November 16, 2020, 05:30:23 pm »
I get this with clang and optimiser doing its thing...

Code: [Select]
        mov     eax, esi
        cmp     edi, esi
        cmovge  eax, edi
Okay, that's where ARM wins.
You no longer compare two words and maybe overwrite one of them, but now you pick one of two words and move it to a third register. This can't be done with CMOV because it's a two argument instruction like most of x86.
The reason those registers are allocated as they are is because of calling convention. The problem would go away if the function is inlined, assuming that code can be generated in such way that one of the inputs enters through the output register.

Good example how minor differences in implementations of almost the same instruction can make or break real world code.

Is this an example of where x86 is, essentially, "wasting time" compared to ARM - by doing unnecessary stuff?
 

Offline Cerebus

  • Super Contributor
  • ***
  • Posts: 10576
  • Country: gb
Re: Apples new M1 microprocessor
« Reply #93 on: November 16, 2020, 05:34:42 pm »
I think I get it - the "footprints in the snow" look the same no matter which order the steps were taken?

Yeah, pretty much like that. I see it as a visual thing, and frankly am far too lazy to do the scribbling of diagrams necessary to do the explanation justice.
Anybody got a syringe I can use to squeeze the magic smoke back into this?
 
The following users thanked this post: SilverSolder

Offline tooki

  • Super Contributor
  • ***
  • Posts: 15924
  • Country: ch
Re: Apples new M1 microprocessor
« Reply #94 on: November 16, 2020, 07:33:52 pm »
It'll be interesting to see what happens and how Apple decide to implement it.

Apple's desktop/laptop market share has held steady for many years now, hovering somewhere between 9.5-10.0%. There hasn't really been any significant growth and anecdotally, I'm seeing and hearing of more people dropping Apple products than there are people switching to them.

They need to stop this cycle of pissing their loyal fan base off, because they are slowly dwindling.

If iMessage were insecure, never mind trivially insecure, there would be rampant abuse of said vulnerabilities. But there isn’t, which certainly lends credence to Apple’s claims. Apple has been making privacy a major pillar of its product advertisement. Don’t you think that is nothing short of an invitation for security researchers to bang on it long and hard? And yet, still no cracks have been found.

I will say this from an actual security researchers point of view: Apple devices are among the easiest to extract data from. I'm not just talking about iMessage, but the entire contents of the device. This historically hasn't been the case but with research came some pretty novel solutions (ones which have existed for years and Apple has barely bothered to do anything about). I suspect the reason why these "vulnerabilities" haven't been patched by Apple is that the tools required are generally reserved for government agencies, law enforcement, and private firms with enough money to spend. It's not something a kid with a laptop can do, so the impact is fairly low to the general population. Secondly, some of these work arounds can't be patched for a variety of reasons, one being that Apple use it themselves for manufacturing and diagnostics.
But Apple has, repeatedly, fixed the vulnerabilities used by those programs. They then find new ones, which Apple then fixes. The classic ongoing cat and mouse.
 

Offline tooki

  • Super Contributor
  • ***
  • Posts: 15924
  • Country: ch
Re: Apples new M1 microprocessor
« Reply #95 on: November 16, 2020, 07:34:31 pm »
If you don't want to believe you never will :horse:

Meanwhile, here's an actual professional cryptographer discussing this exact flaw. A flaw that literal nobody like me expected to exist, obviously. This may give you an idea of how little prestige there is to earn in "researching" it. And of course Apple themselves would never brag about ability to spy on their customers for obvious reasons.
https://www.schneier.com/blog/archives/2015/08/nicholas_weaver_1.html
That’s just more unsubstantiated speculation.
 

Offline magic

  • Super Contributor
  • ***
  • Posts: 8061
  • Country: pl
Re: Apples new M1 microprocessor
« Reply #96 on: November 16, 2020, 07:41:02 pm »
Is this an example of where x86 is, essentially, "wasting time" compared to ARM - by doing unnecessary stuff?
Well, it surely needs to decode one instruction more. Whether it will run slower is hard to say, the first two instructions can be decoded in parallel and executed in parallel and only the CMOV needs to wait for them before executing. The MOV will use some resources while it is in progress, perhaps those could be spent on other instructions, if there are other things that could be done. Thanks to register renaming logic, the MOV probably will not create a pointless copy of ESI value, but it will flip a bunch of transistors in the renaming logic and use some power.

I have almost zero experience with ARM, but off the top of my head there are several differences between the two which may affect performance:
ARM has three operand instructions - the result can be written to a third register which is neither of the source registers
X86 has two operand instructions - the result typically overwrites one of the inputs
X86 can encode memory load and arithmetic in one instruction, ARM needs two instructions for that
X86 has variable length instructions and some common ones are very short
ARM64 doesn't have variable length instructions and is easier to decode
I don't know how real world code density (bytes per actual work done) compares between these architectures, and this too can affect performance because of more or less "work" fitting in instruction cache.

Turing completeness doesn't buy you much. Brainfuck is a Turing complete programming language with only two arithmetic operations - increment and decrement. If you try to make a CPU that way, good luck.

The ultimate proof is in the pudding. Which laptop will be the fastest to compile a million lines of C++ code and to do so without catching on fire? >:D

That’s just more unsubstantiated speculation.
It's very substantiated, by Apple's own documentation. The protocol is easily backdoorable. If you trust Apple, and Orange Fascist's government and its spooks and the secret courts and gag orders, and Apples employees, and their internal procedures to prevent abuse by said employees, and the foreign spies working there, good for you. As far as I'm concerned, it's a gimmick and a marginal improvement over any other system like that. By far the biggest real world threat is passive network interception and plain old TLS takes care of that.
« Last Edit: November 16, 2020, 07:47:59 pm by magic »
 
The following users thanked this post: ve7xen

Offline Twoflower

  • Frequent Contributor
  • **
  • Posts: 745
  • Country: de
Re: Apples new M1 microprocessor
« Reply #97 on: November 16, 2020, 08:06:36 pm »
Is this an example of where x86 is, essentially, "wasting time" compared to ARM - by doing unnecessary stuff?
It's not that simple anymore. Processors do some magic called macro-op fusion. The processor identifies some common instruction groups and fuses them to one internal operation that will be executed in one clock cycle and only uses only a single execution unit. You can probably argue that this causes unnecessary memory fetches (consumes power and takes place in memory/caches) and this additional effort for this translation.
 

Online Berni

  • Super Contributor
  • ***
  • Posts: 5381
  • Country: si
Re: Apples new M1 microprocessor
« Reply #98 on: November 16, 2020, 09:12:59 pm »
Is this an example of where x86 is, essentially, "wasting time" compared to ARM - by doing unnecessary stuff?

Its more that modern CPU architectures constain so many optimizations that you can't simply count instructions.

These CPUs have ways of executing more than 1 instruction per cycle, but it only happens when certain conditions are met. In cases where the next instruction requires the result of the previous instruction this can't happen, but when two or more consecutive instructions are unrelated and all use different subsystems in the CPU they might be executed simultaniusly. Similar for branch prediction and speculative execution, there are areas of the CPU that are already gearing up to execute a branch that is coming up, doing the calculations that only affect internal registers up front so when the branch happens the time taken to flush the pipeline and take the branch is minimised. The branch predictor sometimes gets it wrong too. Compilers tend to know about some of these subtelties and do some work of rearanging non critical instructions to best fill the CPUs pipeline.

So even if one architecutre might use more instructions to do the same thing, its still possible it is faster if the CPU can really pack them into its pipeline efficiently.

You can also see certain workloads be a lot slower on a certain processor just because the code works with too much data to fit into the cache. As soon as you have tight fast loops that try to access too scatered of a memory area that doesn't fit into cache the whole thing slows down by a lot as the core is twidling its fingers and waiting for data to arrive from RAM, while a CPU with a larger cache might be able to fit it in, barely needing to wait at all.
 

Offline brucehoult

  • Super Contributor
  • ***
  • Posts: 6458
  • Country: nz
Re: Apples new M1 microprocessor
« Reply #99 on: November 16, 2020, 09:30:03 pm »
Precisely. It's a 1970s design which has been hacked and hacked and hacked.

ARM however is a 1980s design which has been hacked and hacked and hacked.

Aarch64, which is what Apple's chips run exclusively, is nearly totally different to the 1980s 32 bit ARM design. It is very much a clean sheet post-2000 design, having learned things (good and bad) from PowerPC and Alpha and Itanium. It is much closer to Alpha and to RISC-V than to 1980s ARM.


I was mostly referring to RISC CPUs in general there.

If you find the length of time ago an idea was first proposed / tried some kind of argument against it, I'll just point out that recognizably modern RISC ideas date back to at least 1974 with the IBM 801 and even to 1964 with Seymour Cray's CDC 6600.

RISC-V is also bringing back some other of Cray's brilliant 1970s ideas in a modern setting.
 


Share me

Digg  Facebook  SlashDot  Delicious  Technorati  Twitter  Google  Yahoo
Smf

 

-->