I've never has a 4.1 but I'm familiar with Teensy 4.0.
Exact same MCU, just different pins exposed and more useful stuff on the board (like support for two 8 MByte PSRAM chips).
It's a pretty fast processor for sure, but unfortunately doesn't have NEON (or MVE, which *is* in the more recent M55 and M85).
But 5 bit components do make it easy to do make it easy to do SIMD-in-GPR with 4 or maybe even 5 elements per 32 bits, even with adds that might overflow. Only 3 elements though if you want to do 5x5 multiplies/scaling/blending (and it will have to be multiplying all elements by the same value. not per-element like real SIMD can do)
Yup. In practice, because the 8-bit indexed planes need a look-up anyway, the palette is premultiplied with the corresponding transparency, yielding a 32-bit r10g10b10 color value per pixel, and a 6-bit opacity (32 - transparency). That means that the original background color is first expanded to o5r5o5g5o5b5 pattern –– that's a major chunk of the cycles needed! ––, then multiplied by the opacity, the premultiplied color value added, result shifted down, and ANDed with 31|(31<<10)|(31<<20) in preparation for the additional planes. Further indexed color planes look up the index, the premultiplied palette color value based on the index; and thus require just the multiplication, addition, shift, and AND. Finally, the 5-bit color components are extracted, each looked up in a separate 32-entry 32-bit tables corresponding to GPIO pin configuration. These three are exclusive-OR'd together and the current pin state, the result saved as the new pin state, and written to the DR_TOGGLE register. This allows any output pin order without affecting the other pins, as long as the 15 pins are in the same GPIO bank.
The exact same math works for r10g12b10/r10g11b10 and r5g6b5, too.
With RISC-V
vsmul.vv on 8-bit elements, each element is an
int8_t, and I can use the
c0 + alpha*(c1-c0) form, because
vsmul.vv returns the most significant bits. One subtraction, one multiplication (with right shift), one addition. (I used your machine translated C906 PDF to verify C906 has this.) 24-bit background full color plane would be optimal, as it'd need no expansion, but 16 bits suffices and is denser, thus faster to render to and blit pixmaps to, pixmaps using less memory, and so on; so expanding r5g5b5 to r8b8g8 will likely take more cycles than the calculations do.
To circle back to an FPGA implementation:
For an FPGA, the latency in the operations does not matter as long as it is constant, because this is a pipeline, and the output-to-memory-read latency just does not matter at all. Even a simple barrel multiplier-adder implementing
((64-alpha)*c0 + alpha*c1 + 32) on 6-bit unsigned c0 and c1 and alpha=0..64 (nominally 7-bit), and using only the high 6 bits of the result for the next step, will work absolutely fine. There are three of these in parallel for each plane, so nine operations per pixel, plus of course all that lookup stuff to get the color component and alpha values. It is, essentially, a ridiculously simple fixed arithmetic pipeline. Something like this is included in many Cortex-A7 2D DMA pixel engines, doing exactly these calculations in a separate mini-engine using DMA to read and write the data. And I only need ~ 11 Mpixels/sec throughput (5 bytes per pixel total across planes, with total three color and alpha lookups, one each from a 256-entry LUT). Very, very simple. But, when I started looking at combining this as exposing some expernal memory to a MCU, that's where I found I had no good starting points, resulting in this thread.
128 bit vector registers, so 16 elements in parallel.
Ah, yes. Got a brainfart having read NEON docs at the same time. Three 128-bit vector registers can do 16 pixels in parallel, but the main complexity there is to expand the 16-bit background color to 24-bit. It may end up being more efficient –– doing fractionally more multiplications but wasting less cycles in unpacking the color components –– to use only 75% of each vector register, doing just 4 pixels per 128-bit vector register.
On an MCU, having to strobe the WRITE output pin in the middle of stable data is a yet another pain. I could use a scan-line buffer with GPIO bank toggle data, but either I'd have to waste every second word to toggling just the WRITE pin, or I need careful DMA triggering from a synthesized WRITE pin signal, or a small external logic circuit to generate a 10-20ns delayed pulse (remaining high for 15-25ns) from each rising or falling edge of a dedicated pin, asynchronously.
I like the async edge-to-pulse approach a lot, and would like to integrate that to the carrier board. I haven't investigated that further yet, because I have absolutely nothing that can measure such short pulse durations or time intervals, and would have to rely on trusting simulations and seeing if it works.
Thus far, I've stuffed the WRITE strobe inside the topmost plane scanning, so that it occurs approximately halfway between the data pin toggles. Less than stupidly naïve code does get quite complicated, with this kind of juggling memory accesses, GPIO pin toggling, and ensuring the cpu core pipeline is not stalled for stupid reasons, so this is one of the very rare cases where writing the key code in assembly is definitely warranted. GCC and Clang do make it very easy, allowing extended inline assembly inside C sources, so that one doesn't even have to know the calling conventions; that's how I normally do this. It's also why I don't want to make it public, not yet at least: it's the kind of code that shows the true colors/mettle of a developer, and I'm not mentally ready to be scrutinized at that level. Again. Yet. It's like showing up in public naked.
Considering my very limited budget, but no time limitations, I think a carrier board for Milk-V Duo would be my best bet here. It has 64MiB of fast RAM, a high-speed USB 2.0 that I crave, and two cores, one of which I can dedicate for the framebuffer processing. Then, I can also change the framebuffer design –– limited only by the one C906 core capabilities of compositing it to an external display –– without hardware changes. It is a bit unfortunate, because now I don't have a good motivation to really get into FPGAs.
There is also LuckFox Pico series based on Rockchip RV1103G1/RV1106G3 with a 32-bit Cortex-A7 (1.2 GHz) and a 32-bit RISC-V MCU running at 300 MHz, which has a 2D graphics engine capable of this kind of composition, and 128 MiB of RAM. I prefer the Milk-V Duo (64-bit RISC-V rv64gcv noting the vector engine differences) myself; just pointing out an alternative.
For example 480p@60 Hz video stream
I'm glad someone finally brought this up 
I told everyone in the initial post that I need at most 70 Mbytes/s bandwidth to the FPGA, and described my preference of 320×240 and 480×320 displays. In #12, I reiterated the 40 Mbytes/s to 70 Mbytes/s aggregate bandwidth to the FPGA. In #28, I re-mention the 320×240 70 Hz display modules, then bluntly say "I might want to scale up to 480×320, doubling all of above, but that's the maximum scale I'm interested in. Anything larger, and I'm better off using an embedded Linux stick with OpenGL ES support for controlling the display anyway."
I don't know how to make this any more clear, but I'll try:
I am only interested in display sizes and resolutions up to 480×320 70Hz, i.e. maximum 11 Mpixels/second throughput.
My preferred data formats are one R5G5B5 background plane (16 bits per pixel, 344,064-byte buffer including nonvisible portions; linear access pattern with 22,000,000 bytes/second read rate), plus three overlaid 8-bit indexed color planes (8 bits per pixel, 172,072-byte buffer per plane including nonvisible portions; linear access pattern with 11,000,000 bytes/second read rate each) plus 256-entry A6R6G6B6 LUT palette each, loaded just before the composition of each frame begins, at about 70 frames per second. This computes to around 55 Mbytes/s data read rate in the aggregate. Because of the underlying mechanics, I'd like to do it slightly faster, although I can slightly reduce the frame rate also if needed, so I added a 25% margin. That's where the 70 Mbytes/s aggregate rate comes from.
The existing
un-optimized Cortex-M7 (NXP i.MX RT1062) implementation uses 29 cycles per pixel running at 600 MHz, consuming about 26% of processor time (156M of 600M cycles per second) with 320×240 70Hz displays, and 52% (312M of 600M cycles/sec) with 480×320 70Hz displays. I feel this is wasted, and want to do
more with that processor instead of just running a simple fixed pixel pipeline on this
$22/
$24/
$30 development board. PJRC sells the proprietary bootloader and power sequencing IC, so I can even make my own Teensy-compatible boards. And I'd be happy to use some other MCU as well, as long as it has sufficient RAM, high-speed USB 2.0, and can drive a parallel display controller (~ 22 I/O pins). Of course, I don't want to have to share the only MCU core with the display pipeline either, I just want to separate the pixel pipeline to a simpler slave.
This is not a product, it's my hobby project, albeit with real, non-commercial, use cases helping nontechnical users understand what their Linux appliances are up to, in humorous and visually pleasing manner. (I don't do the graphics myself; I know actual artists who do. Plus I use open source resources like various versions of Tux the penguin.) I do know of several retro-style game projects that would benefit from similar display controller or "graphics card" for a MCU, so I'm hoping to open-source my designs, as soon as I'm comfortable with others assessing my work in public. I'd be very happy to add/modify the planes to tile maps for retro game development, similar to what 90s-era game consoles and arcade games used, for example.
The most pleasing implementation to me would be an FPGA that exposes external memory to a microcontroller, with the FPGA also using a very trivial fixed pixel pipeline with no latency requirements (so even barrel multiplication and addition of the 6-bit color components is absolutely fine, because the display data can lag the framebuffer reads by any amount, it just doesn't matter). That way, I and others could use different microcontrollers, whatever is needed, and even reprogram the pixel pipeline for those retro graphics projects, for example using tile maps or voxel rendering. (I did many such in the nineties myself, and have the best of those squirreled away for future use.)
Although I'm an utter newbie with FPGAs, my background in various programming techniques including descriptive languages (but just no HDLs yet) means I can already
read Verilog, so I know from other projects, and having the existing working MCU implementation, that the pixel pipeline is absolutely no issue here. If I didn't have an existing MCU implementation, I wouldn't know for sure.
What I do not know, is the design points when interposing an FPGA between external memory and a MCU, while the FPGA itself also wants to access that same external memory. I have seen that other projects often struggle with external memory management, although mostly those are with dynamic RAM, not static or pseudo-static or SDRAM, which are what I'd like to use. Most hobbyist ones use development board integrated memories, avoiding that issue altogether. None of my friends I meet in real life do FPGA design, so I'm pretty much in the dark about the
engineering aspects of such designs.
I do not have an exact requirement for the read-write bandwidth allocated to the MCU, because I am in control of that: I'd like to have that in the same ballpark as $1-$4 SRAM/PSRAM chips (in singles at LCSC/Digikey/Mouser, not bulk prices), as that is what these MCUs would use normally anyway. To that end, I imagined that bank-switching, having the MCU access one while the FPGA accesses the other, would be a suitable solution: FPGA would be simpler, basically a cross-point switch, separate from the the pixel pipeline. It's not a design parameter, just an option I can consider and try.
Again, I have a working pixel pipeline implementation on an MCU, but I have nothing regarding the external memory, except one (family) of PSRAM chips that I know works with one of my microcontrollers, and another (much slower) MCU that I'd be happy to use also with an easier parallel async SRAM interface (with multiplexed parallel data and address pins). No designs, no chosen FPGAs, no memory chips selected; only examples of what I already have, so you have an idea what class of hardware I'm looking at. I also have FPGA dev boards to do unit testing and related and unrelated exercises as part of learning HDL, and although I'd like to do that in parallel with the design phase of this project, I think it is separate, normal, expected, and don't have any interest in discussing those learning aspects here.
What kind of help am I looking for, then?
- Pointers on those pitfalls and traps for young players when putting an FPGA between a MCU and external memory.
Clock domains and synchronization I'm already aware of, but otherwise I read that such interfacing is simple, and usually done with FPGA vendor library designs.
- I'd like to know more about the rules on how to choose suitable parts, because I do not have strict requirements yet, and I'd like to balance those requirements with costs, and have the result be achievable to a hobbyist. Also, if that is too much to ask, I'd like to know that too.
- If I wanted the FPGA to look like a 133 MHz QSPI PSRAM using 4 data pins, one clock pin, and a chip select pin, using the same command set that standard APS6404L_3SQR PSRAM supports – simply because that is already extensively tested to work very well with the MCUs with QSPI PSRAM interfaces at 84 MHz clock –, how that affects the design? What kind of 1-8 MiB / 8-64 Mbit RAM would be fast enough for that plus the four aggregate linear burst streams that total to less than 70 Mbytes/second? The 300 MHz+ async SRAM chips that are available at Mouser and elsewhere I consider overkill, not worth their cost.
- If I wanted the FPGA to look like 1 MiB async SRAM (19/20 address pins multiplexed with 8 or 16 data pins, plus the normal selects and strobes) with 8 - 20 ns read and write access cycle times, what are the pitfalls to avoid? Are all, even the cheap FPGAs suitable for this use case?
- I know I can buffer writes and have even slow memory look much faster to the MCU wrt. single operations, as long as the throughput doesn't exceed the external memory capabilities. Fortunately, I expect MCU external memory writes to be vastly more common than reads: the design is such that it eliminates the need for MCU to read the framebuffer when updating graphics. However, not all memory types support a WAIT signal or clock stretching in I2C, the memory telling the host to hold for a bit until the data to be available. I know async SRAM doesn't have that, although some MCUs have an external memory subsystem with Flash/ROM support so they might actually support one even with SRAM. Are there any workarounds to this read timing issue?
- Examples of existing projects, preferably with memory throughputs above 100 Mbytes/s aggregate. It is much easier to emulate someone else poorly as a beginner, than as a beginner emulate a poor implementation but get much better end results.
- Anecdotes, wins, and fails in FPGA-memory or MCU-FPGA-MCU type designs. I like to learn from others mistakes and discoveries; I have no interest or benefit repeating the same mistakes that everyone else does, no matter how traditional and expected that is. It does not matter how in-depth that is, because that just causes me to dig down into the materials I have to understand both the story and the related context. I just tick that way.
Holy hell these posts of mine are long! I regret starting this thread, even though the information has been invaluable to me (Thanks to all, including Smokey!

), because I don't think these posts of mine have enough value per word for others. Sorry.
