Author Topic: Using Claude Code for embedded work  (Read 39353 times)

0 Members and 4 Guests are viewing this topic.

Offline Kjelt

  • Super Contributor
  • ***
  • Posts: 6736
  • Country: nl
Re: Using Claude Code for embedded work
« Reply #300 on: June 16, 2026, 05:24:34 pm »
Just see Claude as a very knowledgeble engineer producing code. You have to judge for yourself whether you trust someone else to work along you.
No you should see Claude for what it is, an AI first generation tool. Nowhere perfect, just starting.
You make it sound like it is a human being  :palm:
 

Online nctnico

  • Super Contributor
  • ***
  • Posts: 30234
  • Country: nl
    • NCT Developments
Re: Using Claude Code for embedded work
« Reply #301 on: June 16, 2026, 06:00:24 pm »
Just see Claude as a very knowledgeble engineer producing code. You have to judge for yourself whether you trust someone else to work along you.
No you should see Claude for what it is, an AI first generation tool. Nowhere perfect, just starting.
Well, the work Claude code is producing for me is on par with that of a talented junior engineer coupled with a large amount of knowledge about standards, design patterns and obtain a sense of what you are trying to achieve quickly. If you have a different opinion, it is probably outdated and I strongly suggest to install Claude code, switch to the opus 4.8 model and throw some real work at it. I'm quite sure you'll be pleasantly surprised. I have been using Claude for a while now for different projects (Python, C, C++, FPGA, making sense of numbers, setting up & analysing automated regression testing, etc) and the results have been outright amazing (and I don't use this word lightly!). I'm not saying you can let Claude do everything by itself; it needs clear instructions but the productivity boost is huge IF you give Claude the right instructions and context.
There are small lies, big lies and then there is what is on the screen of your oscilloscope.
 
The following users thanked this post: hans, bookaboo, Siwastaja, woofy, SpacedCowboy, Dazed_N_Confused

Offline Siwastaja

  • Super Contributor
  • ***
  • Posts: 11241
  • Country: fi
Re: Using Claude Code for embedded work
« Reply #302 on: June 16, 2026, 06:21:23 pm »
Don't you guys think there is "just a little risk" with having some 3rd party tool editing your source code??

Claude does make mistakes...

How is that different from you zipping your project and fetching it back modified, or copy-pasting back and forth? If you carefully read everything you copypaste, you can even more easily carefully diff what modifications it made.

Amount of trust - or amount of human verification and human double checking - is an interesting topic to discuss, but it has absolutely nothing to do with your neutered copy-paste process or lack thereof.
 

Offline Siwastaja

  • Super Contributor
  • ***
  • Posts: 11241
  • Country: fi
Re: Using Claude Code for embedded work
« Reply #303 on: June 16, 2026, 06:51:32 pm »
You make it sound like it is a human being  :palm:

That's really the best way to think about it. The physical structure is not identical, but very similar to human brain; it walks like, it quacks like..... but it isn't.

But attempts to try to treat AI like some special AI thing tend to fail. "Prompt engineering", "context engineering", complicated custom harnesses.

Instead, what works when leading human workers seems to work with leading AI workers:
* Understand the thing you are leading yourself - understand your own requirements
* Explain them not only what you want, but WHY
* Design together by DISCUSSING what needs to be done, why and HOW.
* Prefer positive examples over negative "don'ts"
* Don't waste time at cursing them - keep to the matter
* Don't try to hide information and play games with them; give them ways to work
* Having healthy amount of suspicion - ask why they think what they do; check their work

Everything that fails with human beings fails with AI, too:
* Have no freaking idea what you are building
* As such, give very confusing instructions
* Be very authoritative - design decisions are locked, not up to discussion.
* Don't give them access to your code, only small pieces at the time, without any actual reason (security classification could be such)
* Curse at them
* Having unhealthy amount of suspicion - either none (no checking), or too much (micromanaging)

So treating them like humans is not a bad first instinct. Of course you can optimize the process if you know how the operate: you can just ghost them when done/not needed and they don't feel bad about it. You can get a fresh worker as often as you wish and "fire" it by ghosting it - it doesn't care. And having some understanding on how they operate internally is beneficial. But it's not a top priority thing.
« Last Edit: June 16, 2026, 06:59:57 pm by Siwastaja »
 
The following users thanked this post: nctnico

Offline brucehoult

  • Super Contributor
  • ***
  • Posts: 6464
  • Country: nz
Re: Using Claude Code for embedded work
« Reply #304 on: June 17, 2026, 03:53:32 am »
I know a company of three persons that really benefit from Jira, if only to keep track of the backlog and progress of ongoing tasks and tasks that are on hold due to waiting for some constraint to be fullfilled.
Keeping track of your todo, in progress and on hold tasks/stories could benefit even small companies, heck even individuals that do not have other ways of keeping track of their progress.

To pay big for it or having the multi team interoperability experience is ofcourse overkill in these situations, but the basic functionality it is a really nice tool.

We used to be pretty happy with Bugzilla, which is free and open source and easily hacked-on. It's pretty easy to use it for features and backlog, not only bugs. I see there was a new release in September 2024.
 
The following users thanked this post: Kjelt

Online paulca

  • Super Contributor
  • ***
  • Posts: 6443
  • Country: gb
Re: Using Claude Code for embedded work
« Reply #305 on: June 17, 2026, 07:18:32 am »
We used to be pretty happy with Bugzilla, which is free and open source and easily hacked-on. It's pretty easy to use it for features and backlog, not only bugs. I see there was a new release in September 2024.

Why not wax tablet?
"What could possibly go wrong?"
Current Open Projects:  68000 Self Build computer + OS.
 

Offline Siwastaja

  • Super Contributor
  • ***
  • Posts: 11241
  • Country: fi
Re: Using Claude Code for embedded work
« Reply #306 on: June 17, 2026, 08:07:36 am »
We used to be pretty happy with Bugzilla, which is free and open source and easily hacked-on. It's pretty easy to use it for features and backlog, not only bugs. I see there was a new release in September 2024.

Why not wax tablet?

Why not. Our journey of improvements:
Jira - buggy and annoying to use
linear.app - smooth and non-buggy, very much nicer to use
pen and paper - the ultimate solution, has its own dedicated "monitor" space, never needs login, very intuitive UI: write down the thing you need to remember, strikethrough / cross over it when not relevant anymore.

The key is to avoid postponing work: don't push work into the "write tickets - don't do them - prioritize them in backlog orgies - never do them - only the absolute necessities get done, at last possible time, outside of the ticket system" game. If you want something done, do it right now.

Pen and paper works for short-term context switch memory very well, i.e., dump your mental stack frame when forced to switch tasks. Tickets are too clumsy for that, especially in Jira.

Customer support / bug reports is different, for that you definitely want a ticket system, with the automation of creating the tickets from customer email.
 

Offline Siwastaja

  • Super Contributor
  • ***
  • Posts: 11241
  • Country: fi
Re: Using Claude Code for embedded work
« Reply #307 on: June 17, 2026, 08:28:13 am »
The key is to avoid postponing work: don't push work into the "write tickets - don't do them - prioritize them in backlog orgies - never do them - only the absolute necessities get done, at last possible time, outside of the ticket system" game. If you want something done, do it right now.

Example from 10 minutes ago.

Useful idea -> prompt:
Code: [Select]
❯ FCR calibration view on FCR admin unit management (and only there, keep the Enion-side view intact): highlight offending values in red, bold warning color if:                                           
  * Cha gain outside 0.88 .. 1.08  (check only if DATE CALIBRATED is not invalid)                                                                                                                         
  * Dsch gain outside 0.92 .. 1.12 (ditto)                                                                                                                                                                 
  * abs(ac offs), abs(cha offs), abs(dsch offs) larger than max(400, inverter_rated_power*0.03)                                                                                                           
  * abs(pos_gain) or abs(neg_gain) outside 500..17000                                                                                                                                                     
  * pos_gain more than 20% different compared to the average of pos_gains at other valid different SoC% - and similarly for neg_gain.                                                                     
                                                                                                                                                                                                           
  Additionally: after charge power/discharge powers at every SoC% point, show C-rate with 2 decimals like this: 6480 W (0.61C). Calculate C rate from power / avg(CharCap, DschCap).                       
                                                                                                                                                                                                           
  Red warning highlighting for the C rate:                                                                                                                                                                 
  * If discharge power < 0.40C at SoC 30% or above, colorize.                                                                                                                                             
  * If charge power < 0.45C between SoC 30% - 78%, colorize.                                                                                                                                               
                                                                                                                                                                                                           
  Additionally: Right below the FCR Calibration - hw_id title, show the inverter type and battery type (same as in the unit management page), and the configured battery capacity (Wh) - the stuff in the 
  database. If DschCap (the measured calibration battery capacity) is more than 20% below the configured battery capacity, highlight both the configured value and the measured one in red.               
                                                                                                                                                                                                           
  Feature branch and PR.

That's all. Some minutes later hit merge, and direct into production. And then it's there, doing useful work right away (attachment).

That prompt could go into a ticket, yes, sure. Except I would need to spend more time with it to make sure it's perfect, because the human being who reads the ticket one week later, while I might be somewhere else, would write a comment to the ticket, which I would see 1 week later.

Now an average human programmer would get that perfect ticket maybe 60-70% right. Claude usually gets it 100% right, maybe average is 99% because it sometimes make a mistake.

So, Jira:
Spend 30 minutes on ticket -> discuss with average delay of days to weeks -> get in in backlog -> arrange backlog grooming orgies where you decide not to implement anything because something else is on fire -> never get anything useful to production -> things are in fire because nothing gets done

Direct implementation:
Spend 10 minutes on prompt -> discuss with avarage delay of 2 minutes -> get it in production -> done


 

Offline Kjelt

  • Super Contributor
  • ***
  • Posts: 6736
  • Country: nl
Re: Using Claude Code for embedded work
« Reply #308 on: June 17, 2026, 08:31:18 am »
Customer support / bug reports is different, for that you definitely want a ticket system, with the automation of creating the tickets from customer email.
So you don't have a customer or customer feedback?
Do you have stakeholders? Being colleagues or feature requests or even things that pop up in your mind that are very nice to be implemented some day when time/budget permits? Those are the things you create backlog tickets for with proper prioritization so when you run idle you pick the highest prio remaining ticket.
But for a one man shop it is not suited probably although the paper you wrote your nice to haves on is also probably unfindable.
 

Offline Kjelt

  • Super Contributor
  • ***
  • Posts: 6736
  • Country: nl
Re: Using Claude Code for embedded work
« Reply #309 on: June 17, 2026, 08:37:33 am »
So, Jira:
Spend 30 minutes on ticket -> discuss with average delay of days to weeks -> get in in backlog -> arrange backlog grooming orgies where you decide not to implement anything because something else is on fire -> never get anything useful to production -> things are in fire because nothing gets done

Direct implementation:
Spend 10 minutes on prompt -> discuss with avarage delay of 2 minutes -> get it in production -> done
Jira tickets takes two to three minutes to create, 10 seconds with Outlook email to Jira scripting. One minute to fill additional fields, none in your 1 person company.
The rest of the process is up to you and your company. We discuss the high prio backlog , in progress and waiting (if due time expired) each morning in the DSU.
Customer/stakeholder satisfaction is very high, as required.

In a real company when you have an idea or request and start implementing directly as your example it would be complete chaos. I have experienced this multiple times in the past where SWE picked up the nice creative tasks started working on them while they were low prio nice to have, in the mean time causing frustration with customers, customer support and managers because some hard to find but annoying bugs were not solved.
 

Online paulca

  • Super Contributor
  • ***
  • Posts: 6443
  • Country: gb
Re: Using Claude Code for embedded work
« Reply #310 on: June 17, 2026, 08:55:30 am »
I'm partly jealous.  It's just that scale thing again.  The fractal nature allows patterns to be matched at all levels.  So it sounds like we are talking about the same thing.

I think your approach is actually very close to the OG Agile methodology and you would note it's complete lack of what happens "above the sprint" level.  It has never had opinions on what happens on the longer term or the bigger picture.  That's because it was designed mostly for projects where that fits.  You have a known place, you have incremental changes to make, you accept what you can reasonably get down and plan for just that work.

However, your strategies break down in many situations.  Whether you encounter them or not is a different question.

The customer.  They will give you a 30 minute highlevel "This is what I vision".  You may get some time to process (the larger the company and the larger the deal the longer and more documentation) or they might just ask you, "How much, when can I have it?"

This is where the "only focus on what's next when you get to it" breaks down.  They are not asking how much can you get done in the first 2 or 3 weeks.  They don't care.  They want to know "How much will it cost?" and "When can I have it?"

Even if you dig in and hold onto the "iterate fast" approach you will still need to answer, "How many sprints?"  "How many developers?  How many testers?".  If you answer to the last two is "One because we only have one", it still leaves the first one.  In order to answer it, you will need to have an understanding of what needs done.  Maybe in an embedded project with a few inputs, a few processes and a few outputs you could 'finger in the air' it and get close.  Add a sprint for contingency.

When someone is asking for...  a card payment routing and validation gateway, as a random example, interfacing with a dozen different upstream bank APIs and several dozen downstream 'card machine', ecommerce sites and apps.  "Books"of regulation down to logging requirements and individual container/VM firewalls...  As a minima, the whole thing will need split up into "UI and backend".  The backend will need to be split on "concurrency" and "async" boundaries so it can scale to millions of transactions a minute.  Going through them one at a time is not wise, so you will probably have multiple teams building different parts concurrently.

Good luck answering "How much", "How long" without analysis.  Analysis you can't do in your head, even if you were doing it alone.

How often does this scale occur in the embedded work?  Again it's sort of fractal.  A "dummy" device can be stubbed in for testing so that one team works on one part, whie the other works on another part.  This happens in electronics.  I'm sure it happens in embedded world.  When it's not "cost practical" to force everythng into one MCU/board/component but to span it across several and fork development into different teams.  This is were a proper project planning and task management system might help.

When management start asking, "How much work is on your books for the next 6 months?  Have you got room for a 500k project?"...  that gets awkward if you say, "No idea"
« Last Edit: June 17, 2026, 09:03:07 am by paulca »
"What could possibly go wrong?"
Current Open Projects:  68000 Self Build computer + OS.
 
The following users thanked this post: Kjelt

Offline Siwastaja

  • Super Contributor
  • ***
  • Posts: 11241
  • Country: fi
Re: Using Claude Code for embedded work
« Reply #311 on: June 17, 2026, 09:00:17 am »
So you don't have a customer or customer feedback?

Of course - what we are doing is annoyingly complex so ~3000 customers actually create significant amount of contacts, for which we now have a completely dedicated hired human being - but that's a separate flow. I'm sure no sensible company of any size let customer support tickets flow into their software development flow directly. Like, Outlook->Jira? Sure, if that's a dedicated customer support board, why not. We are actually looking for a good ticket system for customer support cases, and Jira it won't be, that's I'm sure about.

Being colleagues or feature requests or even things that pop up in your mind that are very nice to be implemented some day when time/budget permits?

Yes - and we have over 300 such nice ideas in our Jira backlog.

Ideas are easy to throw around. What to actually implement and when is more difficult. Implementing relatively small but valuable things right away, to give customer value / increase product quality NOW rather than later, greatly reduces the NUMBER of that backlog stuff, and the cost of prioritization / eventual implementation. This is all I am saying.

Truly large ideas that cannot be implemented right now - also need truly bigger specification. They are not that numerous. I don't see any value with a Jira ticket for a bigger idea, anymore. In reality, such ideas are thrown around in Power Points, Excels, Slack, Figma or maybe CAD tools, emails, phone calls and meetings. Sure you can collect those resources and put it in a Jira ticket. Now what do you do next? Ticketize it into 100 sub-tasks and create a timeline for their implementation, and get software engineers work on it for the next 6 months? This is exactly what we tried and tried and tried and which never worked for us. And the big pictures on fire always swamped doing smaller yet semi-important stuff: important but no absolutely critical bugfixes, features that reduce our workload (automation features). This led to feedback loop: our customer support overwhelmed by explaining to customers that no, we have not yet fixed that bug, no we don't yet have this feature everyone is asking for, it's prioritized for later. Wasting 100x more time on unfinished things, rather than just do the things right away.

I mean: this is what a ticket system fundamentally is, by definition: it's a system designed to postpone work. And postponing has a cost. Use with care!

Doing it NOW is CI/CD and I'd say that's a really good pattern (in moderation; of course quality control / testing is important), and agile process where you set up artificial timelines and ticketization processes instead of direct execution is, in that sense, anti-CI/CD.

But maybe Jira and ticket process works better for others. For us it didn't. Maybe we get some sort of ticket systems work again for us, in a positive way; now we are probably overshooting into the opposite direction, being freed from the horror of getting nothing done. We will settle into something sensible.
« Last Edit: June 17, 2026, 09:06:05 am by Siwastaja »
 

Offline Siwastaja

  • Super Contributor
  • ***
  • Posts: 11241
  • Country: fi
Re: Using Claude Code for embedded work
« Reply #312 on: June 17, 2026, 09:42:08 am »
The customer.  They will give you a 30 minute highlevel "This is what I vision".

Note a huge underlying difference that changes the whole discussion:

You are talking from the perspective of a software house; that kind which gets a contract on developing software for others. For you, "customers" are those who order you to build software, for their customers.

I, and I believe, many others on this forum - including peter-h and nctnico - are building software within their own companies; to me and them, "customer" means a totally different thing; it's what you would call maybe end-user.

As a result, nctnico, peter-h or me have quite clear own vision of what we are doing; we are leading our companies (or leading a project/team/subproject within a larger company) and have more freedom of how we achieve the goal. Most importantly, we don't have to care if we follow some software industry practices, because we are not software industry as such. We are offering products/services where software plays a part. We can pick and choose what we want to give a try at; look how peter-h dismisses even stuff like git, and still gets work done, and runs a successful business. Or, how we improve our service level significantly by getting rid of Jira, and using the saved time more productively.

Grumpy "software professionals" can sniff at us, they can ridicule us; we don't care.

Yet, every business was small at start. Every business was lean at start. Complicated processes come later; some are inevitable and beneficial, some are harmful bloat.
 

Online paulca

  • Super Contributor
  • ***
  • Posts: 6443
  • Country: gb
Re: Using Claude Code for embedded work
« Reply #313 on: June 17, 2026, 09:55:39 am »
Complicated processes come later; some are inevitable and beneficial, some are harmful bloat.

I suppose what I am trying to communicate, from a perspective of having sat in many different sizes of company (and customer) is that ...  it scales fast, non-linearly fast.

When you "just add one more thing" to software it does not "add" it multiplies.  It multiplies complexity, communication of complexity, complexity of testing etc.

I tried to get AI to give me a "name" for this, it leaned heavily towards Fred Brooks. (excerpt below).

MCU and MCU containing systems are (or are they not) getting increasingly complex as "Edge computing" is being pushed, as "Cloud intergration" and "IoT" is being pushed.  I am warning that it might seem "incremental" and "easy to manage" but it won't be.  Adapting the appropriate level of best-practices in advance is usually what stops the bloated overkill later in response to allowing it to become unbearable first.

AI:
Quote
Brooks's Law is the most famous formulation: "Adding manpower to a late software project makes it later" — from Fred Brooks's 1975 book The Mythical Man-Month. It captures the superlinear cost of adding people due to communication overhead and onboarding drag.
The underlying scaling dynamic has a few names depending on which angle you're looking at:
Metcalfe's Law (borrowed from networking) describes how communication links scale as n(n−1)/2 — so 5 people have 10 communication channels, but 10 people have 45. This is often cited to explain why team complexity grows non-linearly.
Diseconomies of scale is the economics term for when doubling inputs produces less than double output — software is a textbook case.
Complexity theory in software engineering sometimes uses the term superlinear scaling or combinatorial explosion for how interactions between components, people, and decisions multiply.
There's also Amdahl's Law, which quantifies the ceiling on parallelization gains — relevant when teams try to speed up a project by splitting work.
In organizational theory, this is sometimes called the coordination overhead problem or the n² problem.
The overarching idea — that software resists being managed like a linear industrial process — is what Brooks called the "No Silver Bullet" thesis (his 1986 essay): essential complexity in software cannot be engineered away, only managed.
"What could possibly go wrong?"
Current Open Projects:  68000 Self Build computer + OS.
 

Offline Siwastaja

  • Super Contributor
  • ***
  • Posts: 11241
  • Country: fi
Re: Using Claude Code for embedded work
« Reply #314 on: June 17, 2026, 11:04:18 am »
When you "just add one more thing" to software it does not "add" it multiplies.  It multiplies complexity, communication of complexity, complexity of testing etc.

That's a good thing to recognize. We don't need to accept that as some sort of universal law of physics. Some changes add. Some changes multiply. Some changes devastatingly explode the complexity to third or fourth exponent.

AI might be too eager in doing things too well. For the above example where I wanted suspicious numbers to be in alarm color, it wrote a bunch of tests. Great, feature protected against regressing with tests,  :-+. Then again - what are the actual chances those would accidentally regress? Thresholds are in the code. What could happen is that some poor human being wants to modify those thresholds. Will the test fail now? Test needs to be changed? Or do we need to design configurable thresholds, with threshold configuration UI and database storage?

What a can of worms, for a simple thing. Maybe hard-coded thresholds, without any tests would have been best. So Claude already overachieved from my perspective. From some software professional's perspective, it probably underachieved.

Now, let's say it's a change which just colors alarming values. What are the chances something that simple breaks? Non-zero, but nearly zero, in reality. What is the value of such simple feature? Maybe that means that the poor human being doing the check does 1-2 mistakes less every month. Mistake that means we are paying some penalties to the TSO, and causes those customers to complain, with something very hard to mentally connect to the wrong value in the calibration table, like "my system is selling energy into grid every night at 20-22". So seems like having colorized alarms for values that are off is strong on net-positive side. Better do immediately. I know from experience that average pass time for a ticket like this for our past humans+Jira team (we did have 4 or 5 (I don't even remember which) hired professional programmers at one point) would have been 6 months. With Claude, it was 10 minutes.

I have seen some software professionals fear features. That's understandable because they know how complexity can hit them in the long run. Then again, the value those features add to the table cannot be ignored. Software is all about its features, if it lacks the good features, it's a design mock-up, not an actual project.

Real skill is in being able to implement all important, fancy, even nice-to-have features, while keeping complexity in check. To avoid that high exponent. And this, I believe, is rarely achieved by overengineered architecture. Total ad-hoc mess of course doesn't work either.

And this is a skill AI still isn't fully capable of doing on its own. One of the AI risks. Let it do too large pieces without any supervision, and without any own vision of what the architecture looks like, and you will have an unmaintainable, complex mess.
« Last Edit: June 17, 2026, 11:07:47 am by Siwastaja »
 

Online paulca

  • Super Contributor
  • ***
  • Posts: 6443
  • Country: gb
Re: Using Claude Code for embedded work
« Reply #315 on: June 17, 2026, 11:24:28 am »
Your example of the thresholds is a good one for how it starts.  You write exhaustive tests.  If you change the values, all the test breaks.  You introduce a config mechanism to change the values at runtime or even build time.  However, that itself needs tests.  It adds complexity.  It can be a point of failure itself.

Also, "unit" tests should be unaffected by it.  A unit test should not be coupled to the configuration mechanism, or it can't be tested in isolation.  So the "unit" of code has to "be given" it's lookup table and then queried about it.  Meaning the unit test has to create fabricated config structures and test against those.  However, you also need an integration test which proves the "unit" function at it's intergration points, aka the config system.

Its a can of worms, however, as above, a "pivot" of responsibility doesn't explode out of proportion unless you let it.  Just having the thing which picks the colours get supplied by it's reference data, then you can test it in isolation and make up and reference data you want.  As long as the reference is representative of the real data of course.

"Regression" was your own word.  Tests are often confused with "Prove it works".  That is a tiny fraction of what tests are for.  The biggest being, "After we make a change does everything else still work?"

In a strawman case, someone, maybe even you late one night, used the "color" value to determine the severity in some far flung bit of reporting or logging code.  "Red value scale" say.  Later you change the colour scheme via a theme the user asked for, maybe even for accessibility.  The tests in the Red value scale should fail.  In integration tests at least.  An assumption you made, "Colors are unimportant 'pure' values" was incorrect.  If the tests were written well they should fail.

My personal preference is to add what is needed, which requires that you understand trade offs.  Throwing solutions at undefined problems, "Just in case" is a common practice in the industry which almost has engineering justification - at times - but is gauranteed to generate revenue for software consultancies.  So it is ripe.

AI moves the trade offs in subtle and non-subtle ways.  Usually volume and completeness of tests, documentation, commenting etc.  "Quality costs less effort".

However... quality is not a static target either.  There is always a forward direction to increase it, while not decreasing revenue.
"What could possibly go wrong?"
Current Open Projects:  68000 Self Build computer + OS.
 

Online paulca

  • Super Contributor
  • ***
  • Posts: 6443
  • Country: gb
Re: Using Claude Code for embedded work
« Reply #316 on: June 17, 2026, 11:35:24 am »
Speaking of tests.

How many of you MCU guys actually write "Unit tests"?

I know in "hobbyland" I have rarely bothered, simply because I see the "untestable" nature of the code and my "default", "lazy" answer is... "Nah, not solvable".

However.  I also recall making this determination about a bit of software with international significance.  Once upon a time.  However a very skilled lead engineer proved otherwise.  It took a huge amount of effort, and I do not pretend to understand all the lengths they went to, but he was rambling on about pointer opaquity, compiler firewalling, etc. etc.  Splitting intergation points on tightly defined interfaces so they can be tested in isolation via them.

The furthest I got with MCU code was .. with claude's help.. splitting the layers between "low level register code" and "problem domain logic".  Specifically it was a Parallel Nor FlashROM USB programmer.  The split was the "GPIODriver" and "MemoryTransactions" modularisation.

For testing I wrote a different GPIODriver which did nothing but echo on UART.  My next steps was towards running tests on the code without actually deploying it to the MCU.

EDIT:  One omission of this testing structure is possibly obvious.  It doesn't test "timing".  In big iron world we would use things like strace, nanosecond logging, val-grind, cache grind etc. from critical path analysis, I believe in MCU land the equiv is "signalling states with gpio pins and test points" then attach a scope.

It is a lot of effort, but I think it's worth doing.  Even if not for testing purposes.  When you are in that MemoryTransaction layer, you really don't need to know how "setAddrBit" or "setAddrWord" works.  Does that potentially lead of "missed optimisations?" ... yes.  Can they be fixed later?  Usually.  Are they important?  Depends on the application.
« Last Edit: June 17, 2026, 11:40:44 am by paulca »
"What could possibly go wrong?"
Current Open Projects:  68000 Self Build computer + OS.
 

Offline Siwastaja

  • Super Contributor
  • ***
  • Posts: 11241
  • Country: fi
Re: Using Claude Code for embedded work
« Reply #317 on: June 17, 2026, 12:41:38 pm »
Your example of the thresholds is a good one for how it starts.  You write exhaustive tests.  If you change the values, all the test breaks.  You introduce a config mechanism to change the values at runtime or even build time.  However, that itself needs tests.  It adds complexity.  It can be a point of failure itself.

Also, "unit" tests should be unaffected by it.  A unit test should not be coupled to the configuration mechanism, or it can't be tested in isolation.  So the "unit" of code has to "be given" it's lookup table and then queried about it.  Meaning the unit test has to create fabricated config structures and test against those.  However, you also need an integration test which proves the "unit" function at it's intergration points, aka the config system.

Solution: KISS. Don't overengineer. Don't overtest. Don't prematurely design a "scalable" architecture - those almost always fail. Pay most attention to module interactions; try to keep interfaces clean and simple. This keeps complexity contained.

Human-designed tests were always pain points for us, too. False positives all the time; real regressions passing through. Human employees paid to write tests not understanding the big picture, breaking the functionality, and "fixing" the test to proudly pass that changed behavior.

Software is difficult. There are no silver bullets. But simplicity and not believing into promised silver bullets is the closest one gets.

Ask the AI if all that complexity is needed. Let it propose a smaller solution. So far AI is biasing towards smaller fixes. This is usually a good trait. Except once in a while you truly do need a larger refactor, or create a layer of genericity. Then the human needs to take the lead and make it happen.
« Last Edit: June 17, 2026, 06:46:23 pm by Siwastaja »
 

Offline Kjelt

  • Super Contributor
  • ***
  • Posts: 6736
  • Country: nl
Re: Using Claude Code for embedded work
« Reply #318 on: June 19, 2026, 08:05:40 am »
 :)
 
The following users thanked this post: bookaboo, Siwastaja, 5U4GB

Online nctnico

  • Super Contributor
  • ***
  • Posts: 30234
  • Country: nl
    • NCT Developments
Re: Using Claude Code for embedded work
« Reply #319 on: June 19, 2026, 11:15:02 am »
Your example of the thresholds is a good one for how it starts.  You write exhaustive tests.  If you change the values, all the test breaks.  You introduce a config mechanism to change the values at runtime or even build time.  However, that itself needs tests.  It adds complexity.  It can be a point of failure itself.

Also, "unit" tests should be unaffected by it.  A unit test should not be coupled to the configuration mechanism, or it can't be tested in isolation.  So the "unit" of code has to "be given" it's lookup table and then queried about it.  Meaning the unit test has to create fabricated config structures and test against those.  However, you also need an integration test which proves the "unit" function at it's intergration points, aka the config system.

Solution: KISS. Don't overengineer. Don't overtest. Don't prematurely design a "scalable" architecture - those almost always fail.
I don't agree with not designing a scalable architecture. The problem with designing a good architecture is that it is very hard and time consuming. Though, when done right, a scalable architecture gets a project going quickly AND doesn't get in the way of future expansion.
There are small lies, big lies and then there is what is on the screen of your oscilloscope.
 
The following users thanked this post: bookaboo

Offline tszaboo

  • Super Contributor
  • ***
  • Posts: 9824
  • Country: nl
  • Current job: ATEX product design
Re: Using Claude Code for embedded work
« Reply #320 on: June 19, 2026, 11:36:39 am »
Just checking in.  nctnico, your suggestion about using opus was correct. Opus 4.8 with high effort is what I use now, much better than before, thanks.
 
The following users thanked this post: nctnico, Siwastaja

Offline Siwastaja

  • Super Contributor
  • ***
  • Posts: 11241
  • Country: fi
Re: Using Claude Code for embedded work
« Reply #321 on: June 19, 2026, 03:34:15 pm »
I don't agree with not designing a scalable architecture. The problem with designing a good architecture is that it is very hard and time consuming. Though, when done right, a scalable architecture gets a project going quickly AND doesn't get in the way of future expansion.

Prematurely designed scalable architecture is almost always a mistake. To know what the actual operations need, you need to do some sort of real-world product first. That drives you to see the real pain points, and design a useful architecture which scales into things you need, and doesn't bloat on stuff that's totally irrelevant.

The risk is really that your professionals refuse to work with "hobbyist ad-hoc mess", even when you know that is exactly what you need first to figure everything out.

There is nothing to fear in redesigns and rewrites when they are truly needed. The wheel has been reinvented countless of times, that's why it's better now than the original wooden wheel. When the professionals say "redesign is too costly", they are really saying "we don't have the skills to design this for you", which is a warning sign that you are working with the wrong professionals - choose wisely.

But if a redesign is super expensive, a fresh design of same complexity would have been even more expensive, because it would have lacked all the experience gained from the first iteration. This is the professional's bluff. Their story is always the same: "we should have been done this and that, that would have been easy, but now we can't do anything except an even more expensive super-mega-redesign of two years and millions of €". Bluff.

Software and computer hardware evolves fast, too. Every serious long-term software product has been redesigned and rewritten multiple times during its history, out of necessity. It seems that the more complicated, more "well-designed", scalable product, the faster it spoils.

Claude is like McDonalds of code - you kind of know what you are going to get. Finding right human workers is very difficult and expensive. Some succeed, and that's nice.
« Last Edit: June 19, 2026, 03:50:18 pm by Siwastaja »
 

Offline Dazed_N_Confused

  • Contributor
  • Posts: 19
  • Country: us
Re: Using Claude Code for embedded work
« Reply #322 on: June 19, 2026, 05:13:12 pm »
 In todays Claude Code adventure.   |O
 When ask why codex catches bugs Claude Code doesn't
 
 
« Last Edit: June 19, 2026, 05:14:46 pm by Dazed_N_Confused »
 

Offline Siwastaja

  • Super Contributor
  • ***
  • Posts: 11241
  • Country: fi
Re: Using Claude Code for embedded work
« Reply #323 on: June 19, 2026, 05:56:23 pm »
In todays Claude Code adventure.   |O
 When ask why codex catches bugs Claude Code doesn't
 
  (Attachment Link)

That is the correct answer. You can see it while staying in the same model and harness; just ask Claude to do something, then pop up a fresh session and ask it to check; it will usually disagree on some things, possibly finds real bugs. This is valuable, and the computational cost isn't that high - maybe it took 5M tokens to build a semi-large feature; but it only takes maybe 100-200k to read all that code, reason about it, find a bug or two, and give suggestions.

It's funny how the whole language changes in a fresh context. Just like most human programmers, it's somewhat attached to the code it wrote, and tries to explain things away, it literally says "my code". When you commit that to git and ask a fresh agent to modify the same code, it's not "my code" anymore, then it will be quite critical about it, even if you don't ask it to be critical. It's interesting behavior; maybe more than actual emotion of attachment, it's the goal-driven nature, and when it decides its task is done, it requires a bit of pushing to consider the work unfinished again. A fresh context however, hasn't even started, and it's eager to read and comment on code you even hint looking at.

Using different models probably adds a bit more, because their trainings would have different blind spots, different strengths/weaknesses. But probably difference between Opus 4.8 and GPT 5.5 or Claude Code / Codex harness is secondary; just the fresh context itself is the main factor.

Some try to automate this process, but this is also where human-in-the-loop is beneficial. Maybe you can automate the part where a second agent finds a clear bug, and tells the first one to fix it (or fixes it itself), but usually it's not about clear-cut bugs but rather, finding edge cases / design peculiarities you need human to make decisions.

Practical example: a somewhat complex model predictive control electric boiler controller. I let Claude mostly design the implementation details and implement it. Today fixed three bugs:

1) actual logical error in performance optimization: baseline cost from simulating full plan; plans iterated with partial re-simulation (only after the element which changes) - that partial cost compared to original full-plan cost. This is exactly a type of mistake which AI does not seem to do very often - it's quite good at this kind of logical thinking. Happened nevertheless. Good thing - it found it completely by itself; all I had to do is to persuade it that "yes, it really is broken - create artificial tests and look at their results until you fix it". It fixed it.

2) another similar type of logical error - also found it and fixed it without my help beyond persuading it to find it.

3) most interestingly: and this is something Opus 4.8 truly didn't "grok" itself: predicted hot water usage curve is in "liters of hot water per hour". The model correctly subtracts this much of hot liters from the boiler. But as hot water holds more energy, same liters/hour water usage is modeled as more expensive at boiler temperature of 90degC, compared to 60degC. This drives the algorithm to avoid high temperatures - which reduces spot price arbitrage opportunities ("heat it up to 90degC when electricity is cheap"). Claude says this is correct real-world representation. Many humans would do the same. But what it missed: it missed the mixing valve. If your boiler has 90degC water in it, and you shower with 37degC water, you are running smaller flow rate of that 90degC water. So liter/hour of "hot water" was the wrong thing to begin with; energy would have been correct. I supplied it with the liters/hour number; it was trained with some typical liters/day values it double-checked against; I can't blame it. 99% of hired human programmers would have done the same mistake; my specification lacked this important detail completely, and it's non-obvious. This is exactly the case where AI needs a designer who is really into the thing and understands all the subtleties. This is also something that's easily lost when you write the spec in 2 minutes into Jira ticket. But the nice part is: Claude understands what a mixing valve is; it has all the pieces, it just didn't connect them. Very easy to prompt around; just say "don't forget the mixing valve, with 90degC water less flow is needed, energy is all that matters" - and it fixes the code, runs tests, verifies the result.

But you get only this when you actually sit down and do the actual design work - i.e., chat with the implementer (human or AI). Doesn't work if work is managed by passing tickets back and forth. Then the bug remains and causes slight quality regression no one ever notices, because no one was ever interested in that detail, but managing Jira tickets instead and marking them as done.
« Last Edit: June 19, 2026, 06:04:50 pm by Siwastaja »
 
The following users thanked this post: nctnico, Dazed_N_Confused

Offline Dazed_N_Confused

  • Contributor
  • Posts: 19
  • Country: us
Re: Using Claude Code for embedded work
« Reply #324 on: June 19, 2026, 06:19:49 pm »
In todays Claude Code adventure.   |O
 When ask why codex catches bugs Claude Code doesn't
 
  (Attachment Link)
Using different models probably adds a bit more, because their trainings would have different blind spots, different strengths/weaknesses.

 That's more accurate from my POV. I don't allow CC to do too much before I hit /new after a updated todo list. Still misses lots of issues. New context windows don't normally have big catch flaws improvements for me, but it suppose to save some usage.
Codex promptly finds issues with CC and vise versa.
 Working with AI is becoming a human skill itself.
 


Share me

Digg  Facebook  SlashDot  Delicious  Technorati  Twitter  Google  Yahoo
Smf

 

-->