Author Topic: Hobbyist: FPGA between MCU and external memory?  (Read 30629 times)

0 Members and 8 Guests are viewing this topic.

Offline pcprogrammer

  • Super Contributor
  • ***
  • Posts: 6102
  • Country: nl
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #25 on: February 25, 2025, 08:02:06 am »
As to simulation, mentioned by others, it most likely does not reveal metastability issues unless you simulate very long runs, and even then it might be difficult to spot that it went wrong.

In my case the simulations I ran revealed nothing, but perfectly working logic, that allowed me to tweak some timing issues with passing data from the block memory to the SDRAM. Only when I tried the design on the hardware it revealed problems.

So yes, simulation is very helpful in verifying the logic of the design, and tweaking it when needed, but don't rely on it to be absolutely correct.

Offline brucehoult

  • Super Contributor
  • ***
  • Posts: 6462
  • Country: nz
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #26 on: February 25, 2025, 08:39:17 am »
OP mentioned BL808, but didn't mention CV1800B (available on the Milk-V Duo for $5, which is effectively a DIP breakout board) or SG2000.
Yup.  I already have a SG2002 (Milk-V Duo 256M), its "big brother", too.  CV1800B would do absolutely fine for me, but I don't see any hobbyist designs to compare to.  The schematics (PDF) and docs are available, but I'm not sure if I can pull that off as a hardware design.  I fear I'd be bitten by the power-up sequencing of the four supply voltages it needs (3.3V, 1.8V, 1.35V, and 0.9V).

I suspect my soldering results for 0.35mm pitch QFN would be even worse than they are for 0.8mm pitch BGA. :-[

You can use a Duo as a 2.5mm pitch DIP chip. It's cheaper than many of the MCUs being mentioned -- and I think far cheaper than any FPGA that would do the job, not even counting accompanying MCU.

Quote
In the overlong background, I explained that I already have flat framebuffers in hardware, but would prefer a layered one instead, with one full-color background framebuffer, with two or three indexed-color/paletted overlay framebuffers with full alpha support on top.  I can do this in software on a sufficiently high-powered ARM microcontroller, but the composition is a parallelizable operation best suited for FPGAs and ASICs, not general-purpose processors.

The Duos all do 128 bit vector processing, with over 2.2 GB/s of total DRAM bandwidth available. eg. a 1 MB to 1 MB memcpy() using RVV takes 0.864 ms,  or 1130 MB/s = 2260 MB/s total bandwidth. That's at the default 850 MHz speed of the official buildroot image. I don't think 1 GHz would help it -- the CPU core is already capable of over 3100 MB/s memcpy in cache for transfers from 2k to 16k buffer size.

It's still going to saturate the memory bus with some operation involving reading multiple vector registers of data from different buffers and doing some 8 bit AND / OR / MIN / MAX / blend operation on them.

And don't forget you can control 8 bit or 32 bit pixels with 1 bit per element mask registers, and do AND / OR / XOR etc on masks too.

How fast do you need?

Quote
  I could do this the retro way, using discrete logic and fake ALU chips (MCUs or SRAMs or EEPROMs), all running off a 12-18 MHz bus clock, combining data from 3-4 parallel data buses (retro way using SRAM chips with address coming off a counter).

How wide RAM chips?

4 parallel 8 bit buses at 16 MHz is 64 MB/s, or 256 MB/s with x32 SRAMs. Even the 256 MB/s is only 10% of the 2260 MB/s a CV1800B can do in its 64 MB RAM.
« Last Edit: February 25, 2025, 08:43:00 am by brucehoult »
 
The following users thanked this post: Nominal Animal

Offline Nominal AnimalTopic starter

  • Super Contributor
  • ***
  • Posts: 8349
  • Country: fi
    • My home page and email address
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #27 on: February 25, 2025, 10:15:49 am »
You can use a Duo as a 2.5mm pitch DIP chip.
I didn't think of that, d'oh.

The ALU needed for ARGB processing is a simple one, working with three 6-bit units in parallel, with 6×6=6-bit multiplication operation, and 6-bit addition and subtraction operators.  The planes can also operate in parallel, working on sequential pixels at the same time in a pipeline. 

Essentially, we start with the 15-bit r5g5b5 color, expanding it by copying the MSB to new LSB on each to 18-bit r6g6b6.  I might do some shenanigans with the unused half of the colorspace.  Before generating pixels, at the start of a new frame, each plane reads in a 256-entry A7R6G6B6 palette (25×256 = 6400 bits, probably using 32-bit words or 1024 bytes on the external RAM per plane).  There is plenty of time to do this; it's done at about 14ms intervals (70 Hz or so).  Each plane receives the 18-bit r6g6b6 value from the plane below it in the pipeline, with the full-color one at the bottom, and a stream of 8-bit color index values from the plane buffer..  Each 8-bit value is looked up in the palette, and the two values combined using
    r_new = ((64 - a) * r_0 + a * r_1 + 32) >> 6
    g_new = ((64 - a) * g_0 + a * g_1 + 32) >> 6
    b_new = ((64 - a) * b_0 + a * b_1 + 32) >> 6
and the resulting 18-bit color r6g6b6 passed to the next plane in the pipeline.  At the output, the 18-bit color may be truncated to 15/16 bits, but it's all parallel data with a write strobe to all the controllers I have, running at 6 to 18 MHz clock frequency.  Clock jitter is not a problem, either, the bus is asynchronous.

I wonder if I could use the NPU for this?

With 32-bit general purpose processors, I limit to r5g5b5 (binary 0b00000RRRRR00000GGGGG00000BBBBB) or r5g6b5 (binary 0b00000RRRRR000000GGGGGG00000BBBBB), so it's just two multiplications per plane per pixel, plus additions, shifts, and AND masking to clear the high bits for the next plane.  The palette is similarly expanded, and the result is only compacted for the output pins.  (Teensy doesn't have a contiguous set of I/O pins exposed, so instead I exclusive-OR the result with the previous output word, and three 64-word look-up tables exclusive-OR'd together, saving the result for next, and writing it to the GPIO bank DR_TOGGLE to update the output pins.)

With 6-bit modular arithmetic, special-casing a=0 and a=64, the calculation can be implemented as
    c_new = ((c_0 << 6) + a * (c_1 - c_0) + 32) >> 6
noting that the temporary result is 12 bits.  You don't need to worry about overflow or underflow: as long as you ignore the extra bits, the result will be correct.  (The full ALUs I want need 6+6+6=24 or 6+6+7=25 input bits, and 6 output bits; and I need nine of them.  Or, I can use the expanded form and two 12/13-bit input and 6/7-bit output ones with an adder in Y form.  I never said discrete logic would be trivial, just feasible.  Many FPGAs seem to have hardware single-cycle multipliers, which would be extra nice here.)

I tend to not need the other blitting modes this way.  Even shadows work just fine using RGB blending.  Using La*b* or YCrCb colorspace would give better results, but it's not like I need specific operations per se, just something that looks nice and doesn't limit me in ways I don't like.

My preferred displays are 320×240 with 70 Hz optimal refresh rate.  That corresponds to 10,752,000 bytes of 16-bit color data plus 5,376,000+71,680 = 5,447,680 bytes per plane, per second; say 27,095,040 bytes per second.  Using 6-bit modular arithmetic, I'll need 48,384,000 multiplications and 435,456,000 additions or subtractions per second, which corresponds to about 5,376,000 multiplications per second per arithmetic unit.

On a 32-bit ARM core, it is about 16,128,000 pixel blends per second.  And parallelizes easier than anything I've worked with before.  It's a very clear pipeline.

I might want to scale up to 480×320, doubling all of above, but that's the maximum scale I'm interested in.  Anything larger, and I'm better off using an embedded Linux stick with OpenGL ES support for controlling the display anyway.  (The default non-HDMI TFT LCD for Tang Nano 9k is 800×480.)

I need to check the RISC-V vector instruction set supported, but I do believe Zve32x extension (vmulhu.vx or vmulhu.vv at EEW=8) would work well for me here, using one byte per color component.  In that case, it might make sense to blend one scan line at a time over the three planes in sequence, to keep the data in cache (might need to be written in assembly, with suitable prefetches in place).  I've written MMX/SSE/AVX SIMD assembly code before, both by hand and using GCC vector extensions.

The only thing is that @SiliconWizard is the only one who I've heard has done bare-metal, non-Linux development on Milk-V's.  He posted about this in a thread here.
 

Offline brucehoult

  • Super Contributor
  • ***
  • Posts: 6462
  • Country: nz
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #28 on: February 25, 2025, 11:13:59 am »
I need to check the RISC-V vector instruction set supported, but I do believe Zve32x extension (vmulhu.vx or vmulhu.vv at EEW=8) would work well for me here, using one byte per color component.  In that case, it might make sense to blend one scan line at a time over the three planes in sequence, to keep the data in cache (might need to be written in assembly, with suitable prefetches in place).  I've written MMX/SSE/AVX SIMD assembly code before, both by hand and using GCC vector extensions.

The data and computation rates you're talking about are trivial for a Duo, straight from RAM you don't even need to worry about cache. Just bear in mind that the vector unit is on the "Linux" applications processor, not the MCU core. But you can pass data between them. Or just let uboot SPL to get the hardware up and then simply run your code not the Linux kernel.

Also recall that the vector ISA is XTHeadVector aka RVV draft 0.7.1, which is supported by gcc 14+ vector instrinsics (C source code compatible with RVV 1.0), but also also far far easier to program in asm than AVX etc. For such a simple computation I'd definitely use asm.

Duo's RVV supports 8/16/32/64 bit integers and 32/64 bit FP.

https://github.com/riscvarchive/riscv-v-spec/releases/download/0.7.1/riscv-v-spec-0.7.1.pdf

The optional ediv feature is not implemented (and was dropped from RVV 1.0 in favour of fractional LMUL)
 
The following users thanked this post: Someone, Nominal Animal

Offline Someone

  • Super Contributor
  • ***
  • Posts: 6066
  • Country: au
    • send complaints here
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #29 on: February 25, 2025, 11:42:07 am »
What has been repeated and re-inforced in this thread is "two memories, exclusive access by two masters" but that does not match up with your desired single unchanging background plane with dynamic planes being updated over the top.
No, they are all mutable.  I could show the software implementation of the approach, and how much this reduces memory use and required memory bandwidth if the composition of the planes is done by external hardware.  Simply put, in my use cases and omitting the composition cost, the memory use and number of operations needed per displayed frame is a small fraction of that needed when using a standard flat full-color framebuffer.
Where did I say immutable? If you have two bank switched memories each with exclusive access then any change to any layer must be completed in full, and then applied to the second page/bank (unless it is transient and only appllies for a single step of the animation, which is why I used the unchanging background as the example). This is what you said:

A practical example of this is a small gauge display.  The static part is in the background plane.  The gauge needles and displayed values are on a separate plane on top, so the MCU does not need to modify –– read-then-write –– to change these, just clear and draw the new ones on top.  Because the memory for each plane is separate, clearing the gauge plane does not affect the full-color background at all.
Missing that the background would need to be present in both the banks!

If planes are blocks within a linear space you only have to update the "new"/pending buffer of the plane that is changing, all unchanging content in memory is only in a single place (reducing storage and bandwidth demands).

Just like in software you can build out some stubs/simulations of bits of the system that are large/complex and test the bits you are unsure of under some limited/basic conditions, to check if it the concept is feasible.
Of course.  I've got a "library" (mostly in offline storage) of thousands of such unit test cases in software, examining specific algorithms and math, accumulated throughout the years.
... [describes entire list of specific chips and interfaces already planned]
Sounds like you don't actually want any advice, your funeral.
 

Offline Someone

  • Super Contributor
  • ***
  • Posts: 6066
  • Country: au
    • send complaints here
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #30 on: February 25, 2025, 11:46:52 am »
As to simulation, mentioned by others, it most likely does not reveal metastability issues unless you simulate very long runs, and even then it might be difficult to spot that it went wrong.

In my case the simulations I ran revealed nothing, but perfectly working logic, that allowed me to tweak some timing issues with passing data from the block memory to the SDRAM. Only when I tried the design on the hardware it revealed problems.

So yes, simulation is very helpful in verifying the logic of the design, and tweaking it when needed, but don't rely on it to be absolutely correct.
Which is where the dev kits come in, so you can deploy some slimmed down and/or stubbed out test vehicle to shake out the non-deterministic bits. Without having to worry if it's your undersized power supply or stubs in the DDR RAM routing causing the problem.
 

Offline tszaboo

  • Super Contributor
  • ***
  • Posts: 9823
  • Country: nl
  • Current job: ATEX product design
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #31 on: February 25, 2025, 12:48:42 pm »
If you want to learn FPGAs, than buy an evaluation board
This is not that.  I already have those for that purpose.  This is a specific hardware project with specific issues – FPGA exposing RAM to an MCU, while simultaneously itself accessing that RAM; or using two separate RAM ICs and having the FPGA use one while the MCU uses the other –, and I want to better understand those issues before starting the actual build.
Did you know a good portion of FPGA design are prototyped on dev boards before any hardware is designed/built? While you don't have an architecture planned out (or as above calculated in terms of bandwidths) it is too early to be thinking of what parts are going into the mix.

Just like in software you can build out some stubs/simulations of bits of the system that are large/complex and test the bits you are unsure of under some limited/basic conditions, to check if it the concept is feasible.

What has been repeated and re-inforced in this thread is "two memories, exclusive access by two masters" but that does not match up with your desired single unchanging background plane with dynamic planes being updated over the top. Your description leads towards having multiple buffers for each of the planes that can be swapped in and out per plane, which isn't hard in a single memory address space.

Pointers to existing projects interposing an FPGA between a MCU and RAM (SRAM/SDRAM/PSRAM)?  Especially if some kind of access delegation is used.  This is the part that is most likely to bite me in the butt, so any pointers how to do this right would be welcome.
Having multiple [arbitrary description of separable operations] accessing a single memory is the dominant paradigm of software, and naturally extends to what you are doing. The FPGA/hardware specific counterpoint is synchronous (or close to) streaming which might only be applicable to the pixel output pipeline.

Where is the blockage in connecting [arbitrary write only] master and [arbitrary read only] master to a conventional memory? Ignore the specific interfaces they use for the moment, break it into understandable bits. If that does not make sense then adding multiple readers (and/or writers) will really get you confused.
Exactly, there is no way I would approve the start of a PCB including an FPGA, without seeing first a proof of concept working. And by proof of concept, I mean everything complex demonstrably working. Maybe less sample rate or resolution or communication speed. And especially the most overlooked part, how the entire thing is going to get programmed and verified. I've seen people connect JTAGs together, because "Its all software defined hardware anyway" and "JTAG specification allows this", just to find out you cannot program parts of it, because the toolchain doesn't implement all optional JTAG stuff. The number of pitfalls is very high.
But still my favorite was looking at DDR2 signals, that need to be within 15ps of each other with a 200MHz oscilloscope.
« Last Edit: February 25, 2025, 12:51:05 pm by tszaboo »
 

Offline Nominal AnimalTopic starter

  • Super Contributor
  • ***
  • Posts: 8349
  • Country: fi
    • My home page and email address
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #32 on: February 25, 2025, 06:05:50 pm »
The data and computation rates you're talking about are trivial for a Duo
On Teensy 4.1 (NXP i.MX RT1062 ARM Cortex-M7) using 5 bits per color component, using naïve Arduino C++ code, it is less than 29 cycles per pixel to compose the four-layer framebuffer (likely applicable across ARMv7e-m, but there's plenty of room for optimization here).  For a 320×240 display at 70 Hz, that's 26% of the CPU cycles running at 600 MHz (using 380k of RAM).  And that's with utterly unoptimized code, using plain 32-bit unsigned integer arithmetic, too.  I did say I already have this working, I just want to do it in a better way.

I use completely randomized 320×120-pixel test data, re-randomizing the source data between compositions, the ARM_DWT_CYCCNT cycle counter for the measurements, and running the test for thousands of frames, so this is the baseline to improve on, not an optimistic estimate.

Milk-V Duo is a 64-bit C906, so it won't do any worse, that's for sure.  With 6 bits per color component, vsmul.vv with SEW=8 only needs 8 bits per color component (so 8 color components per 64-bit register; 8 pixels in three registers), so it may do much better.  Expanding the full-color data to a padded form is something that I can for sure improve on, similarly for writing the parallel data to GPIO I/O bank on output; I need to take a careful look at the vector load instructions.

Also recall that the vector ISA is XTHeadVector aka RVV draft 0.7.1, which is supported by gcc 14+ vector instrinsics (C source code compatible with RVV 1.0)
Forgot; thanks!

the background would need to be present in both the banks!
I think me fail English, because to me that is obvious.

The two-bank idea is just a possible workaround, if the framebuffer memory isn't fast enough otherwise.  I'm sure proper designers just upgrade to hardware that is fast enough –– that's what HPC simulator folks do, too, instead of fixing the software instead.  I plan ahead, so I always have alternatives and backup options.

I'm not just jumping into "production" straight away, either.  I'm trying to see what I need to investigate to come up with an FPGA design in the first place, what kind of experiments I need to do to find out my practical requirements, and so on.  This will take a lot of time, but that's fine.

Sounds like you don't actually want any advice, your funeral.
I don't get how you got to that conclusion.  I intended the "I believe I'll start" as my current understanding of the suggested approach with an FPGA, using whatever hardware I already have available as an example, ending with a conclusion that this plan of attack makes sense to me.  Which part of that is rejecting any of the advice suggested thus far?  I'm serious, and want to know.  Or did I just explain it poorly?  (I do fail that way often.)

I do have some hardware already at hand, but I'm not tied to them, although I have plenty time but no money to spend.
 

Offline hgl

  • Regular Contributor
  • *
  • Posts: 110
  • Country: de
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #33 on: February 25, 2025, 06:08:59 pm »
If it is mainly about the memory and not about FPGA you can use this board.

https://de.aliexpress.com/item/1005006147123388.html

Internal Ram 1Mbyte + external Ram 32MByte
The board also has LCD parallel RGB and serial SPI interfaces.

You can read more about this in German here :
https://www.mikrocontroller.net/topic/571964

 
 

Offline asmi

  • Super Contributor
  • ***
  • Posts: 3330
  • Country: ca
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #34 on: February 25, 2025, 06:53:51 pm »
I just want to mention that you'd better make sure your memory has sufficient bandwidth to support your application. Video applications require surprisingly high amount of bandwidth. For example 480p@60 Hz video stream requires 640x480x4x60=70.3125 MiBytes/s of bandwidth for just a single stream, and in your case you need at least 2 streams (1 stream for each layer) + you need to leave some bandwidth for CPU and other processes. Here are some ballpark bandwidth numbers for memories typically used with FPGAs:
Asyncronous SRAM 16 bit / 10 ns2 * (1000 / 10) = 200 MiBytes/s
SDRAM 16 bit / 166 MHz2 * 166 = 332 MiBytes/s
HyperRAM 8 bit / 200 MHz1 * 200 * 2 = 400 MiBytes/s (beware of high initial access latency!!!)
DDR1 16 bit / 200 MHz2 * 200 * 2 = 800 MiBytes/s
DDR2/3/3L 16 bit / 400 MHz2 * 400 * 2 = 1600 MiBytes/s
DDR4 16 bit / 1.2 GHz 2 * 1200 * 2 = 4800 MiBytes/s
 
The following users thanked this post: Smokey

Offline iMo

  • Super Contributor
  • ***
  • Posts: 6898
  • Country: li
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #35 on: February 25, 2025, 07:22:02 pm »
Readers discretion is advised..
 

Offline brucehoult

  • Super Contributor
  • ***
  • Posts: 6462
  • Country: nz
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #36 on: February 26, 2025, 04:23:16 am »
The data and computation rates you're talking about are trivial for a Duo
On Teensy 4.1 (NXP i.MX RT1062 ARM Cortex-M7) using 5 bits per color component, using naïve Arduino C++ code, it is less than 29 cycles per pixel to compose the four-layer framebuffer (likely applicable across ARMv7e-m, but there's plenty of room for optimization here).  For a 320×240 display at 70 Hz, that's 26% of the CPU cycles running at 600 MHz (using 380k of RAM).  And that's with utterly unoptimized code, using plain 32-bit unsigned integer arithmetic, too.  I did say I already have this working, I just want to do it in a better way.

I've never has a 4.1 but I'm familiar with Teensy 4.0. It's a pretty fast processor for sure, but unfortunately doesn't have NEON (or MVE, which *is* in the more recent M55 and M85).

But 5 bit components do make it easy to do make it easy to do SIMD-in-GPR with 4 or maybe even 5 elements per 32 bits, even with adds that might overflow. Only 3 elements though if you want to do 5x5 multiplies/scaling/blending (and it will have to be multiplying all elements by the same value. not per-element like real SIMD can do)

Quote
Milk-V Duo is a 64-bit C906, so it won't do any worse, that's for sure.  With 6 bits per color component, vsmul.vv with SEW=8 only needs 8 bits per color component (so 8 color components per 64-bit register; 8 pixels in three registers)

128 bit vector registers, so 16 elements in parallel.

If you have fewer than 8 variables (I think yes?) then you can pretend you have 512 bit vector registers with LMUL=4. C906 takes 3 cycles per 128 bits for most vector operations, so there is no immediate advantage in using longer LMUL, but it does let you spread the loop overhead and pointer bumps over more data.
« Last Edit: February 26, 2025, 04:27:49 am by brucehoult »
 

Offline Smokey

  • Super Contributor
  • ***
  • Posts: 3879
  • Country: us
  • Not An Expert
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #37 on: February 26, 2025, 08:01:21 am »
I just want to mention that you'd better make sure your memory has sufficient bandwidth to support your application. Video applications require surprisingly high amount of bandwidth. For example 480p@60 Hz video stream requires 640x480x4x60=70.3125 MiBytes/s of bandwidth for just a single stream, and in your case you need at least 2 streams (1 stream for each layer) + you need to leave some bandwidth for CPU and other processes. Here are some ballpark bandwidth numbers for memories typically used with FPGAs:
Asyncronous SRAM 16 bit / 10 ns2 * (1000 / 10) = 200 MiBytes/s
SDRAM 16 bit / 166 MHz2 * 166 = 332 MiBytes/s
HyperRAM 8 bit / 200 MHz1 * 200 * 2 = 400 MiBytes/s (beware of high initial access latency!!!)
DDR1 16 bit / 200 MHz2 * 200 * 2 = 800 MiBytes/s
DDR2/3/3L 16 bit / 400 MHz2 * 400 * 2 = 1600 MiBytes/s
DDR4 16 bit / 1.2 GHz 2 * 1200 * 2 = 4800 MiBytes/s

I'm glad someone finally brought this up :)
 

Offline Nominal AnimalTopic starter

  • Super Contributor
  • ***
  • Posts: 8349
  • Country: fi
    • My home page and email address
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #38 on: February 26, 2025, 09:38:05 am »
I've never has a 4.1 but I'm familiar with Teensy 4.0.
Exact same MCU, just different pins exposed and more useful stuff on the board (like support for two 8 MByte PSRAM chips).

It's a pretty fast processor for sure, but unfortunately doesn't have NEON (or MVE, which *is* in the more recent M55 and M85).

But 5 bit components do make it easy to do make it easy to do SIMD-in-GPR with 4 or maybe even 5 elements per 32 bits, even with adds that might overflow. Only 3 elements though if you want to do 5x5 multiplies/scaling/blending (and it will have to be multiplying all elements by the same value. not per-element like real SIMD can do)

Yup.  In practice, because the 8-bit indexed planes need a look-up anyway, the palette is premultiplied with the corresponding transparency, yielding a 32-bit r10g10b10 color value per pixel, and a 6-bit opacity (32 - transparency).  That means that the original background color is first expanded to o5r5o5g5o5b5 pattern –– that's a major chunk of the cycles needed! ––, then multiplied by the opacity, the premultiplied color value added, result shifted down, and ANDed with 31|(31<<10)|(31<<20) in preparation for the additional planes.  Further indexed color planes look up the index, the premultiplied palette color value based on the index; and thus require just the multiplication, addition, shift, and AND.  Finally, the 5-bit color components are extracted, each looked up in a separate 32-entry 32-bit tables corresponding to GPIO pin configuration.  These three are exclusive-OR'd together and the current pin state, the result saved as the new pin state, and written to the DR_TOGGLE register.  This allows any output pin order without affecting the other pins, as long as the 15 pins are in the same GPIO bank.

The exact same math works for r10g12b10/r10g11b10 and r5g6b5, too.

With RISC-V vsmul.vv on 8-bit elements, each element is an int8_t, and  I can use the c0 + alpha*(c1-c0) form, because vsmul.vv returns the most significant bits.  One subtraction, one multiplication (with right shift), one addition.  (I used your machine translated C906 PDF to verify C906 has this.)  24-bit background full color plane would be optimal, as it'd need no expansion, but 16 bits suffices and is denser, thus faster to render to and blit pixmaps to, pixmaps using less memory, and so on; so expanding r5g5b5 to r8b8g8 will likely take more cycles than the calculations do.

To circle back to an FPGA implementation:

For an FPGA, the latency in the operations does not matter as long as it is constant, because this is a pipeline, and the output-to-memory-read latency just does not matter at all.  Even a simple barrel multiplier-adder implementing ((64-alpha)*c0 + alpha*c1 + 32) on 6-bit unsigned c0 and c1 and alpha=0..64 (nominally 7-bit), and using only the high 6 bits of the result for the next step, will work absolutely fine.  There are three of these in parallel for each plane, so nine operations per pixel, plus of course all that lookup stuff to get the color component and alpha values.  It is, essentially, a ridiculously simple fixed arithmetic pipeline.  Something like this is included in many Cortex-A7 2D DMA pixel engines, doing exactly these calculations in a separate mini-engine using DMA to read and write the data.  And I only need ~ 11 Mpixels/sec throughput (5 bytes per pixel total across planes, with total three color and alpha lookups, one each from a 256-entry LUT).  Very, very simple.  But, when I started looking at combining this as exposing some expernal memory to a MCU, that's where I found I had no good starting points, resulting in this thread.

128 bit vector registers, so 16 elements in parallel.
Ah, yes.  Got a brainfart having read NEON docs at the same time.  Three 128-bit vector registers can do 16 pixels in parallel, but the main complexity there is to expand the 16-bit background color to 24-bit.  It may end up being more efficient –– doing fractionally more multiplications but wasting less cycles in unpacking the color components –– to use only 75% of each vector register, doing just 4 pixels per 128-bit vector register.

On an MCU, having to strobe the WRITE output pin in the middle of stable data is a yet another pain.  I could use a scan-line buffer with GPIO bank toggle data, but either I'd have to waste every second word to toggling just the WRITE pin, or I need careful DMA triggering from a synthesized WRITE pin signal, or a small external logic circuit to generate a 10-20ns delayed pulse (remaining high for 15-25ns) from each rising or falling edge of a dedicated pin, asynchronously.

I like the async edge-to-pulse approach a lot, and would like to integrate that to the carrier board.  I haven't investigated that further yet, because I have absolutely nothing that can measure such short pulse durations or time intervals, and would have to rely on trusting simulations and seeing if it works.

Thus far, I've stuffed the WRITE strobe inside the topmost plane scanning, so that it occurs approximately halfway between the data pin toggles.  Less than stupidly naïve code does get quite complicated, with this kind of juggling memory accesses, GPIO pin toggling, and ensuring the cpu core pipeline is not stalled for stupid reasons, so this is one of the very rare cases where writing the key code in assembly is definitely warranted.  GCC and Clang do make it very easy, allowing extended inline assembly inside C sources, so that one doesn't even have to know the calling conventions; that's how I normally do this.  It's also why I don't want to make it public, not yet at least: it's the kind of code that shows the true colors/mettle of a developer, and I'm not mentally ready to be scrutinized at that level.  Again.  Yet.  It's like showing up in public naked.

Considering my very limited budget, but no time limitations, I think a carrier board for Milk-V Duo would be my best bet here.  It has 64MiB of fast RAM, a high-speed USB 2.0 that I crave, and two cores, one of which I can dedicate for the framebuffer processing.  Then, I can also change the framebuffer design –– limited only by the one C906 core capabilities of compositing it to an external display –– without hardware changes.  It is a bit unfortunate, because now I don't have a good motivation to really get into FPGAs.

There is also LuckFox Pico series based on Rockchip RV1103G1/RV1106G3 with a 32-bit Cortex-A7 (1.2 GHz) and a 32-bit RISC-V MCU running at 300 MHz, which has a 2D graphics engine capable of this kind of composition, and 128 MiB of RAM.  I prefer the Milk-V Duo (64-bit RISC-V rv64gcv noting the vector engine differences) myself; just pointing out an alternative.

For example 480p@60 Hz video stream
I'm glad someone finally brought this up :)
I told everyone in the initial post that I need at most 70 Mbytes/s bandwidth to the FPGA, and described my preference of 320×240 and 480×320 displays.  In #12, I reiterated the 40 Mbytes/s to 70 Mbytes/s aggregate bandwidth to the FPGA.  In #28, I re-mention the 320×240 70 Hz display modules, then bluntly say "I might want to scale up to 480×320, doubling all of above, but that's the maximum scale I'm interested in.  Anything larger, and I'm better off using an embedded Linux stick with OpenGL ES support for controlling the display anyway."

I don't know how to make this any more clear, but I'll try:

I am only interested in display sizes and resolutions up to 480×320 70Hz, i.e. maximum 11 Mpixels/second throughput.

My preferred data formats are one R5G5B5 background plane (16 bits per pixel, 344,064-byte buffer including nonvisible portions; linear access pattern with 22,000,000 bytes/second read rate), plus three overlaid 8-bit indexed color planes (8 bits per pixel, 172,072-byte buffer per plane including nonvisible portions; linear access pattern with 11,000,000 bytes/second read rate each) plus 256-entry A6R6G6B6 LUT palette each, loaded just before the composition of each frame begins, at about 70 frames per second.  This computes to around 55 Mbytes/s data read rate in the aggregate.  Because of the underlying mechanics, I'd like to do it slightly faster, although I can slightly reduce the frame rate also if needed, so I added a 25% margin.  That's where the 70 Mbytes/s aggregate rate comes from.

The existing un-optimized Cortex-M7 (NXP i.MX RT1062) implementation uses 29 cycles per pixel running at 600 MHz, consuming about 26% of processor time (156M of 600M cycles per second) with 320×240 70Hz displays, and 52% (312M of 600M cycles/sec) with 480×320 70Hz displays.  I feel this is wasted, and want to do more with that processor instead of just running a simple fixed pixel pipeline on this $22/$24/$30 development board.  PJRC sells the proprietary bootloader and power sequencing IC, so I can even make my own Teensy-compatible boards.  And I'd be happy to use some other MCU as well, as long as it has sufficient RAM, high-speed USB 2.0, and can drive a parallel display controller (~ 22 I/O pins).  Of course, I don't want to have to share the only MCU core with the display pipeline either, I just want to separate the pixel pipeline to a simpler slave.

This is not a product, it's my hobby project, albeit with real, non-commercial, use cases helping nontechnical users understand what their Linux appliances are up to, in humorous and visually pleasing manner.  (I don't do the graphics myself; I know actual artists who do.  Plus I use open source resources like various versions of Tux the penguin.)  I do know of several retro-style game projects that would benefit from similar display controller or "graphics card" for a MCU, so I'm hoping to open-source my designs, as soon as I'm comfortable with others assessing my work in public.  I'd be very happy to add/modify the planes to tile maps for retro game development, similar to what 90s-era game consoles and arcade games used, for example.

The most pleasing implementation to me would be an FPGA that exposes external memory to a microcontroller, with the FPGA also using a very trivial fixed pixel pipeline with no latency requirements (so even barrel multiplication and addition of the 6-bit color components is absolutely fine, because the display data can lag the framebuffer reads by any amount, it just doesn't matter).  That way, I and others could use different microcontrollers, whatever is needed, and even reprogram the pixel pipeline for those retro graphics projects, for example using tile maps or voxel rendering. (I did many such in the nineties myself, and have the best of those squirreled away for future use.)

Although I'm an utter newbie with FPGAs, my background in various programming techniques including descriptive languages (but just no HDLs yet) means I can already read Verilog, so I know from other projects, and having the existing working MCU implementation, that the pixel pipeline is absolutely no issue here.  If I didn't have an existing MCU implementation, I wouldn't know for sure.

What I do not know, is the design points when interposing an FPGA between external memory and a MCU, while the FPGA itself also wants to access that same external memory.  I have seen that other projects often struggle with external memory management, although mostly those are with dynamic RAM, not static or pseudo-static or SDRAM, which are what I'd like to use.  Most hobbyist ones use development board integrated memories, avoiding that issue altogether.  None of my friends I meet in real life do FPGA design, so I'm pretty much in the dark about the engineering aspects of such designs.

I do not have an exact requirement for the read-write bandwidth allocated to the MCU, because I am in control of that: I'd like to have that in the same ballpark as $1-$4 SRAM/PSRAM chips (in singles at LCSC/Digikey/Mouser, not bulk prices), as that is what these MCUs would use normally anyway.  To that end, I imagined that bank-switching, having the MCU access one while the FPGA accesses the other, would be a suitable solution: FPGA would be simpler, basically a cross-point switch, separate from the the pixel pipeline.  It's not a design parameter, just an option I can consider and try.

Again, I have a working pixel pipeline implementation on an MCU, but I have nothing regarding the external memory, except one (family) of PSRAM chips that I know works with one of my microcontrollers, and another (much slower) MCU that I'd be happy to use also with an easier parallel async SRAM interface (with multiplexed parallel data and address pins).  No designs, no chosen FPGAs, no memory chips selected; only examples of what I already have, so you have an idea what class of hardware I'm looking at.  I also have FPGA dev boards to do unit testing and related and unrelated exercises as part of learning HDL, and although I'd like to do that in parallel with the design phase of this project, I think it is separate, normal, expected, and don't have any interest in discussing those learning aspects here.

What kind of help am I looking for, then?
  • Pointers on those pitfalls and traps for young players when putting an FPGA between a MCU and external memory.

    Clock domains and synchronization I'm already aware of, but otherwise I read that such interfacing is simple, and usually done with FPGA vendor library designs.
     
  • I'd like to know more about the rules on how to choose suitable parts, because I do not have strict requirements yet, and I'd like to balance those requirements with costs, and have the result be achievable to a hobbyist.  Also, if that is too much to ask, I'd like to know that too.
     
  • If I wanted the FPGA to look like a 133 MHz QSPI PSRAM using 4 data pins, one clock pin, and a chip select pin, using the same command set that standard APS6404L_3SQR PSRAM supports – simply because that is already extensively tested to work very well with the MCUs with QSPI PSRAM interfaces at 84 MHz clock –, how that affects the design?  What kind of 1-8 MiB / 8-64 Mbit RAM would be fast enough for that plus the four aggregate linear burst streams that total to less than 70 Mbytes/second?  The 300 MHz+ async SRAM chips that are available at Mouser and elsewhere I consider overkill, not worth their cost.
  • If I wanted the FPGA to look like 1 MiB async SRAM (19/20 address pins multiplexed with 8 or 16 data pins, plus the normal selects and strobes) with 8 - 20 ns read and write access cycle times, what are the pitfalls to avoid?  Are all, even the cheap FPGAs suitable for this use case?
     
  • I know I can buffer writes and have even slow memory look much faster to the MCU wrt. single operations, as long as the throughput doesn't exceed the external memory capabilities.  Fortunately, I expect MCU external memory writes to be vastly more common than reads: the design is such that it eliminates the need for MCU to read the framebuffer when updating graphics.  However, not all memory types support a WAIT signal or clock stretching in I2C, the memory telling the host to hold for a bit until the data to be available.  I know async SRAM doesn't have that, although some MCUs have an external memory subsystem with Flash/ROM support so they might actually support one even with SRAM.  Are there any workarounds to this read timing issue?
     
  • Examples of existing projects, preferably with memory throughputs above 100 Mbytes/s aggregate.  It is much easier to emulate someone else poorly as a beginner, than as a beginner emulate a poor implementation but get much better end results.
     
  • Anecdotes, wins, and fails in FPGA-memory or MCU-FPGA-MCU type designs.  I like to learn from others mistakes and discoveries; I have no interest or benefit repeating the same mistakes that everyone else does, no matter how traditional and expected that is.  It does not matter how in-depth that is, because that just causes me to dig down into the materials I have to understand both the story and the related context.  I just tick that way.



Holy hell these posts of mine are long!  I regret starting this thread, even though the information has been invaluable to me (Thanks to all, including Smokey!  :-+), because I don't think these posts of mine have enough value per word for others.  Sorry.  :-[
« Last Edit: February 26, 2025, 09:48:24 am by Nominal Animal »
 

Offline brucehoult

  • Super Contributor
  • ***
  • Posts: 6462
  • Country: nz
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #39 on: February 26, 2025, 09:38:27 am »
I just want to mention that you'd better make sure your memory has sufficient bandwidth to support your application. Video applications require surprisingly high amount of bandwidth. For example 480p@60 Hz video stream requires 640x480x4x60=70.3125 MiBytes/s of bandwidth for just a single stream, and in your case you need at least 2 streams (1 stream for each layer) + you need to leave some bandwidth for CPU and other processes. Here are some ballpark bandwidth numbers for memories typically used with FPGAs:
Asyncronous SRAM 16 bit / 10 ns2 * (1000 / 10) = 200 MiBytes/s
SDRAM 16 bit / 166 MHz2 * 166 = 332 MiBytes/s
HyperRAM 8 bit / 200 MHz1 * 200 * 2 = 400 MiBytes/s (beware of high initial access latency!!!)
DDR1 16 bit / 200 MHz2 * 200 * 2 = 800 MiBytes/s
DDR2/3/3L 16 bit / 400 MHz2 * 400 * 2 = 1600 MiBytes/s
DDR4 16 bit / 1.2 GHz 2 * 1200 * 2 = 4800 MiBytes/s

I'm glad someone finally brought this up :)

As I already pointed out, the 64 MB of DDR2 in-package in the CV1800B SoC (on e.g. the $3 Milk-V Duo board) does 2260 MB/s total bandwidth on an actual, real world, memcpy() test, not some theoretical speed-of-light.
« Last Edit: February 26, 2025, 09:46:22 am by brucehoult »
 

Offline Nominal AnimalTopic starter

  • Super Contributor
  • ***
  • Posts: 8349
  • Country: fi
    • My home page and email address
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #40 on: February 26, 2025, 09:57:27 am »
And I've repeatedly told that the calculations show the FPGA only needs 70 Mbytes/second or less, in four separate linear cacheable/bufferable streams.  It is the unpredictable MCU memory accesses, and how to cater for both of these in the same design, preferably keeping to the performance one can expect from say APM6404L QSPI PSRAM, that makes me scratch my head.  Writes I can buffer, but those pesky unpredictable reads!

CV1800B just has a dual-core rv64gcv processor, and I don't mind "wasting" the second core for the composition, because the rest of what I want to do with the same MCU, graphics update stuff and such, doesn't need nor profit from two cores.  It is still wasting a 64-bit core for what a simple FPGA can do, but at these prices, it just Makes Sense.
« Last Edit: February 26, 2025, 09:59:55 am by Nominal Animal »
 

Offline brucehoult

  • Super Contributor
  • ***
  • Posts: 6462
  • Country: nz
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #41 on: February 26, 2025, 10:41:23 am »
And I've repeatedly told that the calculations show the FPGA only needs 70 Mbytes/second or less, in four separate linear cacheable/bufferable streams.

Right, which makes me wonder why people act like this is a difficult data rate to handle.

Quote
It is still wasting a 64-bit core for what a simple FPGA can do, but at these prices, it just Makes Sense.

Have you got an estimate for just how simple an FPGA, and the price?
 

Offline Nominal AnimalTopic starter

  • Super Contributor
  • ***
  • Posts: 8349
  • Country: fi
    • My home page and email address
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #42 on: February 26, 2025, 01:08:05 pm »
Have you got an estimate for just how simple an FPGA, and the price?
No, because just the framebuffer without external memory is useless to me.  And I like to attack the hardest parts first, leaving the design finalization to when I have a firm grasp on the details.

It's responding to the MCU read and write requests efficiently and fast enough, while also keeping those data pipelines to the compositor full, that make me scratch my head.  I'd like to know more about it in general, before attacking it.  Thus, this thread.

I suppose that I really should have implemented the framebuffer on say Tang Nano, just to find out how much resources the implementation consumes.  I didn't, because it is a problem I've already "solved".  And, with the huge uncertainty I have whether I can do the external memory and MCU interfaces, it just wasn't interesting enough as-is.

There are also FPGA-specific details, like what kind of lookup units and clock frequency it supports, that determines exactly how the core arithmetic, ((64-a)×c0 + a×c1 + 32)>>6 = (64×c0 + 32 + a×(c1-c0))>>6 = c0 + ((a×(c1 - c0) + 32)>>6) is best implemented.  If one has 4096×6 block RAM, then that can do all the multiplications I want, one each in a single block RAM access.  If it can do single-cycle lookups at 100 MHz, that only leaves the RAM lookups and additions and subtractions to do, with their latency irrelevant.  However at 100 MHz, I have 9 cycles per pixel, which means even a barrel multiplier (a×b = a0*(b) + a1*(b<<1) + a2*(b<<2) + a3*(b<<3) + a4*(b<<4) + a5*(b<<5)) suffices, if I can have 9 of them in parallel.  And some FPGAs have fixed single-cycle multiplication units built-in; as long as each term is 6 bits or more, and includes 7 most significant bits, they work for me, without any compromises needed.

Even the memory access for the framebuffer is relatively simple, because it is linear data, and easily cached.  The only "sequencing" in that is that I want to read the palettes (256 entries per plane, three planes; each entry specifying 6-bit r, g, b, and 6/7-bit opacity/alpha, 0..64) first just before starting the four framebuffer data streams in parallel (staggered) –– in suitable blocks, depending on how much bits I have to cache this data in a shift-register-like fashion.  So, obviously some kind of memory arbitrator is needed.  Where to find more about those?

I do understand why selecting the FPGA first is so important, but I'd like to know more about how to do that selection well.  It seems like many just pick a FPGA, and then work within its capabilities.  I don't like that approach at all; I don't thrive in walled gardens, no matter how nice.
« Last Edit: February 26, 2025, 01:14:11 pm by Nominal Animal »
 

Online xvr

  • Frequent Contributor
  • **
  • Posts: 916
  • Country: ie
    • LinkedIn
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #43 on: February 26, 2025, 02:23:51 pm »
Quote
I do understand why selecting the FPGA first is so important
No. More important to select memory configuration first. FPGA is second.

Let's consider different configurations:

1. 2 separate RAM for framebuffers (as you suggest as possible configuration). This is will not work out (unless you are ready to fill all 4 planes for new framebuffer before each new frame gone to LCD).

2. PSRAM for framebuffer + PSRAM interface between FPGA and MCU will not work out. PSRAM has no possibly to postpone transaction, so MCU interface can't handle PSRAM access when physical PSRAM will be busy talking to LCD framebuffer circuitry. Started LCD access can't be paused to handle MCU one.

3. SRAM for framebuffer + PSRAM interface to MCU could work out, if SRAM will have enough bandwidth (70 MB/s + PSRAM bandwith for MCU interface). This approach required carefully planned timing sharing between MCU and framebuffer interfaces. Especially for MCU part, because it required not only a bandwidth but limited access latency.
 
The following users thanked this post: Nominal Animal

Offline dolbeau

  • Regular Contributor
  • *
  • Posts: 100
  • Country: fr
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #44 on: February 26, 2025, 02:54:49 pm »
Pointers to existing projects interposing an FPGA between a MCU and RAM (SRAM/SDRAM/PSRAM)?

Not an MCU per se, but I have given access to FPGA DDR3 memory to a MC68030, MC68040, or SuperSPARC CPU for two use cases:

(a) the memory is used as a framebuffer, with the FPGA playing the part of the display device connected to a screen. The CPU needs to access the framebuffer to display things (read and write, not just write). The FPGA is connected either to the SBus (SPARCstation 20), NuBus ('030 Mac, '040 Mac), or directly to the CPU bus ('030 Mac, '040 Mac)

(b) the memory is used as actual memory (usable by the operating system), in this case only for the MC68030-based Macintosh IIsi, directly on the CPU bus (using whatever memory is not used for the framebuffer)

For (a) the memory is also used from inside the FPGA, as there is a VexRiscV core with some micro-code for display acceleration.

The main issue for (b) is that the (read) latency is higher than for the "real" memory in the system. In my case, the '030 bus is synchronous (20 MHz), the request then has to cross to the internal bus of the FPGA SoC (100 MHz system clock Wishbone, no relationship to the '030 clock) to access the memory controller, and then all the way back. Write are just immediately acknowledged and put in a FIFO (and read do not bypass at all, they have to wait for the FIFO to be empty, which doesn't help read latency either, could be improved) so are pretty quick. For read, I ended up adding support for the MC68030's burst mode, so each request brings back 128-bits (the width of the memory controller) which are sent back to the CPU in 4 bus cycles after the long wait to get it from the DDR3. It helps.

Other than that, mostly smooth sailing hardware-wise as the memory controller I used (litedram) support multiple ports and does the arbitration for me.
 
The following users thanked this post: Nominal Animal

Offline pcprogrammer

  • Super Contributor
  • ***
  • Posts: 6102
  • Country: nl
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #45 on: February 26, 2025, 02:56:46 pm »
To make a complete picture every aspect is important. If you want to connect two 32 bit wide SRAM chips, you need the select a FPGA with enough IO (pins) to connect these and still have enough left to connect the LCD panel and the MCU.

Using SDRAM uses less IO pins, but is more difficult to work with. The MCU interface is also an issue. Will it be a parallel bus or serial and if so with how many lanes.

Choosing the FPGA also depends on the other selections. Do you need high speed serial connections and with what type of signals, like lvds or lvcmos at what ever voltage. The FPGA has to support all these choices. Most of the modern ones will, but you have to look at the specifications to be sure, and do the IO banks with their supplied voltages allow for all the connections.

For instance, when you choose DRAM that needs 1.8V on the pins, you have to setup IO banks with a supply of 1.8V. This means that all the pins of these banks only work with 1.8V. If the MCU needs 3.3V, you have to setup an IO bank with a supply of 3.3V, etc.

Then the number of logic elements needed, how many block ram, how many dsp blocks, etc.

So to make a choice it is a bit of the chicken and the egg story.

As for loading data into the LCD all that is needed to make it work properly with the memory is using a double buffered line buffer. The memory controller writes to line A at high speed and the LCD reads from line B at a lower speed. Once the LCD is done with the line, swap the buffers and trigger another read from the memory.

Mixing of planes can be done before the LCD. Just setup four line buffers that are filled by the memory controller and mix the lines before they are send to the LCD. As long as the mixing can be done in a single pixel time it should not be a problem.

My advice is to design a block diagram with the separate parts of the system and try to work out the needed data streams and data manipulations. Then try to write things up in a HDL and run it through a simulator. This will show if your logic is valid. Then choose an IDE for an FPGA. Don't look at yosis, but an IDE from a vendor like Quartus from Altera(Intel), Vivado from Xilinx(AMD) or Gowin FPGA designer. Select a FPGA and run your design. The IDE will tell if it will fit or not. When making the final choice leave some room for expansion or problem solving.

Or just start to play with your Tang Nano 9K to see if you can get a frame buffer to work with the PSRAM and a LCD. Then expand things to suit your needs. That is how I learn, but your miles may vary.

On a side note, it turns out that the SDRAM in the Tang Nano 20K is rather sensitive to timing issues. At 166MHz using a 180 degree phase shifted clock gave problems, that disappeared when making it the same phase, but not al the time, depending on the place and route. Have to test with 45 or 90 degree phase differences to see if I can get it stable.

Edit: fixed some typos
« Last Edit: February 26, 2025, 03:01:42 pm by pcprogrammer »
 

Offline Nominal AnimalTopic starter

  • Super Contributor
  • ***
  • Posts: 8349
  • Country: fi
    • My home page and email address
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #46 on: February 26, 2025, 09:04:32 pm »
No. More important to select memory configuration first.
Right; that makes much more sense.

2 separate RAM for framebuffers (as you suggest as possible configuration) [..] will not work out (unless you are ready to fill all 4 planes for new framebuffer before each new frame gone to LCD).
It does work just fine; it just works in a very different way that those used to single linear framebuffers are used to.

Let's say we have three separate RAM chips, each with their own buffer contents A (=Abg+Aovl1+Aovl2+Aovl3), B, C.  If the microcontroller does nothing but just repeatedly tell the FPGA to switch to the next buffer at the next retrace, the physical display will show a three-frame animation.

Yes, it can feel wonky that when you update buffer contents, the existing contents are then from three frames back, not from the previous frame.  It is not a problem.  This triple buffering is a very traditional computer graphics method when the frame update takes longer than the display update does, but you avoid both tearing effect and having to wait for the next retrace before updating the display contents for the next frame.  Nowadays, this is still used in best video players, because it allows the processor to exceed single frame update duration without glitching the output, plus negative audio delay for proper sync without having to decouple the audio and video input stream buffering, as the triple buffering adds at least one frame (20ms or so) of extra video delay.

I am extremely familiar with this technique.  For Damage Buffers (a bitmap describing which areas of the buffer need to be updated due to previous modifications), they are maintained for each buffer, not just the current frame.  It is all easily managed by the MCU or whatever is scribbling to the display buffers; the techniques are well known, albeit probably esoteric now that we have hardware that does not usually need such tricks.

PSRAM for framebuffer + PSRAM interface between FPGA and MCU will not work out.
This is what I assumed, based on the datasheet of the PSRAM I'm already using.  (It'd need to be magically faster PSRAM than what the MCU uses now, and that's not realistic.)

SRAM for framebuffer + PSRAM interface to MCU could work out, if SRAM will have enough bandwidth (70 MB/s + PSRAM bandwith for MCU interface). This approach required carefully planned timing sharing between MCU and framebuffer interfaces. Especially for MCU part, because it required not only a bandwidth but limited access latency.
Righto.

Let's assume the MCU uses QSPI at 85.7 MHz (=600/7), with the CLK always running (so the FPGA can PLL on to that, and the MCU-FPGA-RAM uses that as the only clock domain). Each cycle takes 11.7ns. The address takes 6 cycles, with data bytes taking two cycles each; or a byte every 23ns.  I can control the number of cycles in between the address and the first half of the first data byte, but let's say four cycles for reads, none for writes.

It seems to me that if I use async SRAM with 10ns read and write cycle times, half the cycles satisfy the MCU needs.  That doesn't leave me with sufficient framebuffer bandwidth, though.

If I up the QSPI clock to 120 MHz (=600/5), each QSPI clock cycle takes 8.3ns.  Let's assume I use two ISSI IS61WV5128EDBLL-10TLI (PDF) in parallel (10ns read and write cycles at 3.3V), with 19 common address outputs, 16 data I/Os, common chip enables and output enables but separate write enables (=4 outputs), for 39 I/Os for SRAM.  Then, every other cycle (16.7ns intervals, plenty of time) the SRAM provides 16 data bits, and half the accesses satisfy the MCU, leaving 60 Mbytes/s for the framebuffer.  This would be a bit tight, but still suffice.

I could then use 120/7 = 17.1 MHz or 120/8 = 15 MHz clock for the display bus, keeping everything synchronized to the one single clock domain.  The FPGA then needs 39+6+23 = 68 I/O pins at minimum.

Using four of those SRAM chips in parallel would require 18 additional I/O pins, for a total of 86 I/Os minimum, but then there'd be plenty of data bandwidth to the SRAM.  (A pair of 4Mb 10ns async SRAM is half the cost of a 8Mb 10ns async SRAM.)

Does the above estimation look sane, along similar lines you'd do?  (I do not mean to imply I've decided on the above; I'm just trying to show how I understand the advice thus far, and the design choices I believe I would make if I made them now based on that advice.)

Tests that I'd do before locking in any design choices, would include a FPGA PSRAM interface at 120 MHz continuously running clock to verify MCU-FPGA read access by synthesizing the data from the address bits (reordering and XOR'ing); experimental synchronous time multiplexing SRAM arbitrator, to understand the complexity and resources needed; and the actual SRAM ICs connected to an FPGA to test data integrity and timing.  For example, in Intel/Altera MAX 10 datasheet, I saw that they do provide implementations for various RAM types, except async SRAM is "do your own".
« Last Edit: February 26, 2025, 09:06:53 pm by Nominal Animal »
 

Offline asmi

  • Super Contributor
  • ***
  • Posts: 3330
  • Country: ca
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #47 on: February 26, 2025, 09:17:37 pm »
If I up the QSPI clock to 120 MHz (=600/5), each QSPI clock cycle takes 8.3ns.  Let's assume I use two ISSI IS61WV5128EDBLL-10TLI (PDF) in parallel (10ns read and write cycles at 3.3V), with 19 common address outputs, 16 data I/Os, common chip enables and output enables but separate write enables (=4 outputs), for 39 I/Os for SRAM.  Then, every other cycle (16.7ns intervals, plenty of time) the SRAM provides 16 data bits, and half the accesses satisfy the MCU, leaving 60 Mbytes/s for the framebuffer.  This would be a bit tight, but still suffice.
I don't think it will suffice capacity-wise. 320*240*4=300 KiBytes, you will need two of those for layers, and three - for triple-buffering (that will require even more bandwidth BTW as now you simultaneously read 2 layer streams, then write result to a buffer, and have one more stream which actually output video signal to a display, which comes to 4 streams total), while two 512Kx8 devices will only give you 1 MiB.
« Last Edit: February 26, 2025, 09:19:43 pm by asmi »
 

Offline Nominal AnimalTopic starter

  • Super Contributor
  • ***
  • Posts: 8349
  • Country: fi
    • My home page and email address
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #48 on: February 26, 2025, 09:37:08 pm »
If I up the QSPI clock to 120 MHz (=600/5), each QSPI clock cycle takes 8.3ns.  Let's assume I use two ISSI IS61WV5128EDBLL-10TLI (PDF) in parallel (10ns read and write cycles at 3.3V), with 19 common address outputs, 16 data I/Os, common chip enables and output enables but separate write enables (=4 outputs), for 39 I/Os for SRAM.  Then, every other cycle (16.7ns intervals, plenty of time) the SRAM provides 16 data bits, and half the accesses satisfy the MCU, leaving 60 Mbytes/s for the framebuffer.  This would be a bit tight, but still suffice.
I don't think it will suffice capacity-wise. 320*240*4=300 KiBytes, you will need two of those for layers, and three - for triple-buffering (that will require even more bandwidth BTW as now you simultaneously read 2 layer streams, then write result to a buffer, and have one more stream which actually output video signal to a display, which comes to 4 streams total), while two 512Kx8 devices will only give you 1 MiB.
Once again, I'm not doing traditional framebuffers.  I've explained this already above.  I'm using one 16-bit full-color background layer, with three indexed-color overlays.  Maximum size I'm interested in is 320×480 pixels, so 320×480×(2+1+1+1) = 768,000 bytes per frame; × 70 frames/second yields 53,760,000 bytes.  The palette data is 3×256×3 bytes × 70 frames/second = 161,280 bytes, for an aggregate total of 53,921,280 bytes per second RAM aggregate bandwidth.

The desired total framebuffer area is 352×512 pixels; the extra 32 pixels along each axis are not displayed.
352×512×(2+1+1+1) + 3×256×3 = 903,423 bytes.
512k×(8+8) bits = 512k×2 bytes = 1024k = 1,048,576 bytes.
Because 1,048,576 > 903,423, it will suffice.

Why you keep ignoring all the data I post?  I fully understand if TL;DR.  But, ignoring posts and just responding based on your own assumptions and preconceptions about what you believe is being discussed without reading any of it is quite annoying.
 

Online xvr

  • Frequent Contributor
  • **
  • Posts: 916
  • Country: ie
    • LinkedIn
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #49 on: February 26, 2025, 10:15:21 pm »
Quote
It does work just fine; it just works in a very different way that those used to single linear framebuffers are used to.
Yes, it will work in this way (but quite complicated from software point of view)

Quote
Does the above estimation look sane, along similar lines you'd do?
Yes. But it will require a lot of FPGA pins to connect SRAM chips. I afraid that such FPGA will be available only in BGA package. It could be problem for hobby project.
 
 


Share me

Digg  Facebook  SlashDot  Delicious  Technorati  Twitter  Google  Yahoo
Smf

 

-->