Yeah, now models like e.g. GLM-5.3-Flash can be run on hardware that is within the reach of SMEs (~£10-20k) then we know that the big guys are going to struggle to compete. These models are nearly as capable as Claude Opus 4.7-4.8.
Is there
any public example demonstrating this "nearly as capable" factor or is it just wishful thinking?
Further, why would the original authors whose work was copied "struggle" to compete? By definition, those who develop are always ahead of those who copy. The only factor where copies win would be price.
There are public demonstrations of these full size versions, like the GLM-5.3 (full size version, non-quantized, post-trained by illicit use of Anthropic/OpenAI/Gemini/Grok infra) being capable of roughly the Opus 4.5-4.7-level performance, but that does not automagically generalize into order of magnitude smaller distilled versions. If that was true, why the big players don't serve such smaller versions of their models, why Google's small model results embedded in the search results are still utter crap, and why everyone still uses the big models to do any serious work?
So "nearly" as good, nah, jury is still out on that. If they did work, people would use them to do the same tasks, so why doesn't that happen? If something sounds too good to be true, then why not consider the possibility it likely isn't? Why speculate, go do some real work with these small models? So far it's all
