Author Topic: Hobbyist: FPGA between MCU and external memory?  (Read 31638 times)

0 Members and 3 Guests are viewing this topic.

Offline Nominal AnimalTopic starter

  • Super Contributor
  • ***
  • Posts: 8349
  • Country: fi
    • My home page and email address
Hobbyist: FPGA between MCU and external memory?
« on: February 24, 2025, 08:17:20 am »
The question is the difficulty level for a hobbyist to use cheap FPGAs to expose SRAM or SDRAM or PSRAM to a microcontroller.  This would be my first real FPGA project, and I'd like to avoid BGA, too.  If you don't already know, I'm a lightweight very-limited-budget hobbyist in electronics, total newbie with FPGAs, but have decades of programming experience (including some low-level, bare metal stuff) in about a dozen programming languages.

I have a dual purpose for this thread: one is to perhaps get advice on starting this hobby project, and the other (more important to me) is to discuss similar use cases and projects others might be working on.  I don't know how to simplify this starting post without losing context, so sincere apologies for those who hate the walls of text!

Edited to clarify:

The core of the question is about the technical and especially design aspects of putting an FPGA between a memory IC and a microcontroller.  Or, equivalently, an FPGA with built-in memory between two identical microcontrollers; the MCU or MCUs see only ordinary external memory.  In my case, I don't necessarily need access arbitration, because I can use two RAM ICs with separate buses, with each MCU accessing their own one, never the same RAM, as long as I can swap which one they access.  I believe this is a relatively unusual design, with the FPGA-uses-own-RAM being the common one, with only rarer devices like oscilloscopes and other data acquisition equipment sharing memory this way, so I want to learn about the important design-affecting issues and approaches up front.

The rest of this post is details on what and why, but the above is the part that gives me pause.

Background

I like to interface small 1" - 4" display modules to Linux-based appliances and single-board computers for use as an external information display for nontechnical humans.  High-fidelity IPS panels up to 800×480 (my preference is 320×240) are quite cheap.  I can do this using various microcontrollers with USB connection to the host, with the host software only controlling the display, not generating the image.  The issue is that typical microcontrollers only expose a single framebuffer, and that's just not very effective.  The displays have their own controllers and easy digital interfaces, so I find the idea of using an FPGA as a programmable framebuffer extremely interesting.  I'm only a hobbyist, so I want to keep the costs down and everything hand-solderable, so the optimum would be to use external SRAM/SDRAM/PSRAM as a framebuffer exposed to a microcontroller to draw onto, with the FPGA updating the display module transparently on the background.

In particular, I like the idea of having 3-4 completely separate planes: a background one with full 16/18-bit color, plus 2-3 various indexed/paletted planes drawn on top.  Similar stuff was used in arcade games and home game consoles, back when memory was scarcer.  (Even today, with our stupidly powerful GPUs even in phones, the actual displayed image is composited from separate "sub-frames" in graphics memory.)

A practical example of this is a small gauge display.  The static part is in the background plane.  The gauge needles and displayed values are on a separate plane on top, so the MCU does not need to modify –– read-then-write –– to change these, just clear and draw the new ones on top.  Because the memory for each plane is separate, clearing the gauge plane does not affect the full-color background at all.  Plus, the MCU does much less work.

I often use capacitive touch to control what is displayed, so yet another plane can be used for the touch interaction; I like rings or circular arcs extending outwards from the touch point like a ripple to indicate touch registration (separate from touch action on the graphics).  As a separate plane/framebuffer in memory, the ripples can just be copied to the framebuffer by the MCU without any calculation, or any effect on what is being done to the other planes.

If there is a fourth plane available, it can be used in between, for additional overlays and transients.  The indexed color mode I'm most interested in these is one byte per pixel, with each of the 256 possible values freely chosen from a full-color palette (262k colors) with a full transparency control.  For creating antialiased text and polygons, using 3 or 4 low bits for transparency makes things easier for the MCU, but precise and clean graphics and antialiased text.

It would also be fun to play with the kind of visual models used in old game consoles.  All my hobby stuff like this is CC0-1.0/Public Domain, so if I get this working, it'd also be available for other hobbyists for use in old-style portable pixel art and retro gaming projects, too.

As an end result, I'm hoping to end up with a display carrier board with an FPC connector to the display modules using ILI9341/9488/ST7789 etc. controllers, containing the MCU, external RAM, FPGA, USB connector, and perhaps a microSD socket for graphics asset storage.  I currently only use Linux and open source tools, so I'd prefer to avoid proprietary software; and I don't have funds for developer kits.

Hardware:

Among the many microcontrollers from NXP, STmicroelectronics, and others, I've taken a liking to WCH CH32V307VCT6: it has a high-speed USB 2.0 (480 Mbit/s) and an external parallel static memory interface (FSMC) interface for up to 16 MBytes (but with data and address bits multiplexed on the same pins).  It's cheap (<5€ in singles), programmable using open source tools in Linux, simple enough for a hobbyist like myself to put on a custom board, and although not terribly powerful, definitely powerful enough for this.

As to the memory itself, I've been thinking about using two SRAM modules, something like IS61/64WV51216EEBLL 512k×16 SRAM (PDF datasheet at Mouser), with one "connected" to the MCU while the FPGA reads the other for the framebuffer generation.  As I've not done any real projects with FPGAs yet, this is purely to keep the two separated, so that I don't have to work out a bulletproof sharing logic between the FPGA framebuffer read access and MCU read/write access.  Having access only to the part of the framebuffers not currently being displayed is absolutely fine for me, even if it means extra cost: this is hobbyist one-offs, not production stuff.  I do believe CH32V307VCT6 memory interface has a WAIT input I might use to interleave FPGA and MCU accesses, but I know how risky that kind of complexity is in a hobby project, so I'm aiming to keep it simple, even if it means extra hardware.  The downside is that using two in parallel would take quite a lot of pins: about 40 per SRAM module, another 24 to the MCU, and another 24 or so for the display module.

Even faster SDRAM might make things easier.  I do need very fast random access for the MCU (20 ns write cycle time or less).  While the FPGA needs similar throughput (I'd like to get about 70 Mbytes/s), it's almost all just linear scanning.  With fast enough SRAM, the FPGA could probably use fixed time division multiplexing approach, using "access slots" for the MCU, the FPGA frame buffer, and SDRAM refresh.  I know nothing about how to control SDRAM or any of the other RAM types; SRAM thus far seems the simplest to implement.  I'm very open to suggestions, though.

As to the FPGA, I have no idea.  I prefer to avoid BGA, which limits my selection, especially as I'd like to have > 100 I/O pins.  I'd like to have fast 6×6=12-bit multiplication, because optimally I'd do 24 of those in parallel in the output stage (per pixel), with a throughput of about 10 Mpixels/second.  Otherwise the logic is limited to marshalling the RAM accesses and RAM lookup.  Many FPGAs seem to support different voltage levels on different banks.  The MCU and display I/O will be 2.8V–3.3V.  Again, this is my first real FPGA project, so any suggestions would be appreciated.

As to the display modules, I have a few 320×240 and 480×320 IPS TFTs from BuyDisplay/EastRising.  I personally also like the smaller 1"-2" ones, but for nontechnical people, the display needs to be 2.4" – 4" to be practically useful.

Existing software implementation:

I already have experimentally implemented the image generation stuff on a computer and on an MCU, so it's just the hardware that gives me pause. I have considered using a dual-core microcontroller, with one core dedicated for the display generation and updates, but most of the cheap ones just do not have enough RAM for multiple framebuffers/planes, or they support full-speed USB 2.0 only (12 Mbit/s, not high-speed 480 Mbit/s) –– for example, Raspberry Pi RP2350 and Bouffalo Labs BL808 ––, and I feel those are limits I don't want.
(I also have a largeish collection of pixel tricks and effects dating back to nineties.  I particularly like retro arcade and puzzle games with orthographic projection views, like in Gauntlet and similar, and have some tricks for those if anyone is interested.)

The software is all experimental code, not published yet, but I like this "new" approach (well, going back to roots of computer graphics, retro game consoles and the like, except with better hardware) A LOT.  So, this is not a question of "is this possible", but "how could I do this in hardware", and have fun at the same time.  And hobbyist-level stuff only!  I don't care if someone else develops this further into a product, but I definitely will not.  If the end result works, I'll put it all out in OSHW, likely CC0-1.0/Public Domain.

(CH32V307VCT6 has only 64kBytes of RAM, which does not suffice even for a small full-color framebuffer.  The FSMC interface is not pin-compatible with any SRAM chips I've found, because of the way it multiplexes the address and data pins.  I could maybe design a circuit for the needed glue logic, or just use the host computer/appliance and just stream the framebuffer changes.. but I don't want to.)

My favourite microcontroller, Teensy 4.1 (based on NXP i.MX RT1062), has a high-speed USB 2.0 (that even with just plain old USB serial in Linux can sustain 25 Mbytes/s in one direction), lots of RAM (1 MBytes) plus PSRAM support (up to 16 MBytes), and has a Cortex-M7 running at 600 MHz so it has no issues doing this in software.  It's pinout isn't optimal for driving these external display modules, and although PJRC sells the proprietary bootloader chip (which also handles power sequencing on startup) so I could design my own board with a more betterer pinout, it is 0.8mm pitch BGA, and I still fail at BGAs.  In other words, I can do this on relatively cheap hardware already; it's just that an FPGA-based reprogrammable hardware approach exposing the framebuffer/graphics resources as external memory would be even more useful.
Assuming it's not expensive, of course.

Advice needed:

Suggestions for the FPGA?  I know of Lattice iCE40, and have a Sipeed Tang Nano 1k and 9k development boards.  So, a complete beginner.  I'd like to spin my own boards, and avoid BGA.  The displays are in the 2.4"-4" range, so board space is not an issue; I'll use JLCPCB or similar for the boards.  I'm not interested in expensive developer boards and kits, because I have a tiny budget.

Pointers to existing projects interposing an FPGA between a MCU and RAM (SRAM/SDRAM/PSRAM)?  Especially if some kind of access delegation is used.  This is the part that is most likely to bite me in the butt, so any pointers how to do this right would be welcome.  I only need a couple of megabytes of RAM total, so no need to go into dynamic RAM and their intricacies: I'd like to keep this simple.

Advice, articles, projects, or books showing how to interface SRAM/SDRAM to an FPGA correctly?  I like async SRAM because it is simple.  SDRAM requiring periodic refresh cycles scares me a bit, especially because the MCU will not be aware of those.  I can control the async timing on the MCU, but I'd rather keep it as fast as possible to make maximum use.  (The CH32V307 maximum clock frequency is 144 MHz, but I'd be happy with 20ns read and write cycles to the managed external memory.)

Note the tiny budget, though, and this being hobbyist only, and that I'm a Linux-only user right now: I do not have any Windows licenses, and relying on Wine feels clunky, so any Windows-only toolchains are not possible for me.  I'm hoping I can find something supported by Yosys, nextpnr, and similar open source projects, although I'm happy with free vendor tools too, as long as they run in Linux.
« Last Edit: February 24, 2025, 02:04:36 pm by Nominal Animal »
 
The following users thanked this post: Tantratron, incf

Offline pcprogrammer

  • Super Contributor
  • ***
  • Posts: 6110
  • Country: nl
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #1 on: February 24, 2025, 08:40:43 am »
Hi Nominal Animal,

sounds like an interesting project.

On the FPGA side the Gowin GW2AR-18 ones are very interesting. Unfortunately getting them in Europe seems to be a bit difficult and to my surprise LCSC does not sell them either. But the eLQFP144 pin ones cost about 10 dollars in the USA. https://www.edgeelectronics.com/fpga/gw2ar-lv18lq144c8-i7/

One benefit is that you can get versions that have 8MB of SDRAM in the chip with a 32 bit wide databus and 166MHz max clock frequency. A read burst of 256 words takes ~1.55us at this clock rate.

Hardware design is not that complex for these devices. Just take a look at the crappy schematics of the Sipeed Tang Nano boards.

Since you are new to FPGA design and usage of SDRAM, you have some learning to do. I also just recently started with this to make a video conversion system. https://www.eevblog.com/forum/projects/please-recommend-a-vga-to-parallel-lcd-board-or-ic/msg5828029/#msg5828029

SDRAM is not that complex. You can write your own controller for it to suit your needs.

Be aware that programming for a FPGA is very different to writing software. I suggest taking a look at verilog tutorials (I find this easier than VHDL). https://www.asic-world.com/verilog/veritut.html

I'm happy to share my recently gained knowledge with you.

Success,

Peter

Online SiliconWizard

  • Super Contributor
  • ***
  • Posts: 17799
  • Country: fr
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #2 on: February 24, 2025, 08:53:44 am »
About the CH32V307, note two things:
- You can actually use up to 128KB RAM (it's a configuration and will decrease the amount of Flash available for code by 64KB).
- It has an interface for 8-bit and (IIRC?) 16-bit PSRAM. Those chips are easy to find, but possibly not as easy to assemble as they are usually BGA (unless you can find some in another package). They are the same as the small quad spi PSRAM, but with a wider synchronous data bus (so faster access). The CH32V307 can use these, from what I had read in the datasheet. The interface is usually called "Hyperbus". One example: https://www.digikey.com/en/products/detail/infineon-technologies/S27KL0642DPBHI020/11611394

Now the CH32V307 possibly doesn't have enough processing power to do what you want by itself even with that much RAM.

Lattice iCE40 FPGAs look like a good and easily available option. As for RAM, you can go for either SRAM or PSRAM or SDRAM. I think PSRAM will be your cheapest bet here for a given amount of memory, and implementing access in HDL is not too difficult. Note though that anything that's not SRAM will have bursty I/O, meaning that either you the delay for initiating bursts of consecutive data blocks is OK in your design, or you'll need to implement some kind of cache to make it a bit more transparent. But either way, you'll have to account for possible delays when accessing data.
 
The following users thanked this post: Nominal Animal

Offline Nominal AnimalTopic starter

  • Super Contributor
  • ***
  • Posts: 8349
  • Country: fi
    • My home page and email address
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #3 on: February 24, 2025, 12:06:43 pm »
I'm very aware of the difference between Verilog/VHDL/etc. and application software development, but it is an important point to note.  I too see learners often struggling with their existing notions conflicting in a new domain, for example moving from application development to embedded development using freestanding C and/or C++.  Even moreso when changing paradigms (imperative vs. functional and so on).

I do have a Sipeed Tang Nano 9K with a HDMI connector for playing with full-size displays, and a Tang Nano 1K for controlling small TFTs.  Their USB is limited to full-speed only, which I don't want to be limited to.  (I guess I've been overpampered with Teensy 4.x.)

Note: The core of this project is very similar to using an FPGA to provide dual-port access to external (standard) RAM.  It's just that one "port" is the FPGA itself reading the needed framebuffer data out in a very predictable fashion.

CH32V307 isn't grunty enough to implement my framebuffer idea in software; even on Teensies the use of PSRAM must be limited to keep the data bandwidth high enough with the standard PSRAM chips (currently APS6404L_3SQR at 84 MHz).  It's just that it's a simple, cheap MCU with high-speed USB and a simple external RAM interface I think I could manage; compare to e.g. NXP i.MX RT1060 series and their power up sequencing and so on.  The display data generation itself is parallelizable to any degree (down to one pixel per task), with per-task work just memory lookups and basic arithmetic on 4-15 bit integers, which I believe maps well to FPGAs.

(There are ARM Cortex M7 and M33 MCUs from Microchip and ST and NXP I could use instead, but thus far there have been too many for me to find all suitable ones – with high-speed USB and an external RAM interface I could hijack for this, enough internal RAM, and non-BGA packages – that I as a hobbyist could use. Microchip SAMS70, for example.)

Teensy 4.1 itself uses QSPI (SDR) to talk with the external PSRAM.  It'd be really nice to interface the display FPGA with that, especially because there is support for two PSRAM chips (same bus, different chip selects), and it has both RAM and grunt aplenty – and the SDIO micro-SD slot built-in.  Making my own Teensy-type board for displays is in my todo list, but thus far, I've failed with BGA ICs.  I do have now a small hot plate and hot air and a couple of different cheap irons, but no oven.  I really need cheap BGAs to train with, I guess.

As long as the data transfer between the FPGA and the MCU is within electrical specs, I don't mind the MCU framebuffer access occasionally being delayed some clock cycles.  I'm much more worried about getting the approach technically correct, and not relying on "tricks" or picking parts that exceed specs, and getting sufficient write throughput for the MCU.  That's also why I don't just cobble something together and call it done.

You might remember my working out the TFT backlight driving some years back here, because I detest flicker, and some of the displays I have use LEDs in parallel so are not suitable for the cheap flicker-free backlight series-LED drivers.  Combining ultrasonic PWM and (peak) current control is what I like.  "Cheap but over-engineered", I guess!
« Last Edit: February 24, 2025, 12:14:46 pm by Nominal Animal »
 

Offline pcprogrammer

  • Super Contributor
  • ***
  • Posts: 6110
  • Country: nl
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #4 on: February 24, 2025, 12:32:17 pm »
A QSPI interface between a MCU and a FPGA is simple. Lots of examples around for SPI, and adapting for 4 bits in parallel is simply adding shift registers for them. Depending on the external wires a 100MHz should not be a problem. No idea what the Teensy supports.

An important thing in the FPGA side of this is to be aware of metastability. To make what you want having different clock domains is unavoidable, because the LCD most likely needs a different clock speed to push the data to it, the memory might run on a much higher clock and the bus with the MCU will also have it's own clock rate.

To cope with crossing these domains internal dual port memory can be used to handle the data, but for handshaking it is needed to synchronize the signals to the clock of the domain they enter.

This was a problem I encountered in my project. I had setup the double buffered dual ports and triggered the state machines with handshake signals, but without synchronization it locked up almost instantly in the real hardware. In simulation it looked perfect though. Now that I added the needed synchronization to the signals, it runs rock steady. See: https://cdrdv2-public.intel.com/650346/wp-01082-quartus-ii-metastability.pdf

The Tang Nano 9K can also control your smaller LCD's and allows you to play with what I wrote above. For the PSRAM used in the 9K the controlling is different than for the SDRAM in the 20K, but the principle will be the same.

What I find to be a downer with the Tang Nano boards is that they used the 88 pin FPGA to keep it narrow enough, but this does not provide a lot of useful pins on the headers. Sure there are a lot of pins, but when you use the LCD connector maybe half of them can't be used anymore. Also you can't use and the LCD and the HDMI at the same time. They are also a bit pricey in comparison to what you get, at least to my taste.


Offline tszaboo

  • Super Contributor
  • ***
  • Posts: 9824
  • Country: nl
  • Current job: ATEX product design
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #5 on: February 24, 2025, 12:39:52 pm »
Here is an AI summary for that wall-of-text:
Quote
You’re a budget-conscious hobbyist with programming experience, new to FPGAs, aiming to build a low-cost, hand-solderable display carrier board using a WCH CH32V307VCT6 MCU and an FPGA to manage external SRAM (or SDRAM/PSRAM) as a framebuffer for 1"-4" display modules, supporting 3-4 separate planes for retro-style graphics. The FPGA would update the display while the MCU draws, avoiding BGA packaging and complex RAM-sharing logic, with a target of >100 I/O pins and fast pixel processing. You’ve prototyped the software, prefer open-source Linux tools, and plan to release the project as CC0-1.0/Public Domain for hobbyist use. You seek FPGA suggestions, similar projects, and simple SRAM/SDRAM interfacing resources.
My suggestion is this:
If you want to learn FPGAs, than buy an evaluation board, and connect it to another evaluation board. Because when your hadware doesn't work for whatever reason, you will be stuck, and no way to understand or correct why it doesn't work.
 
The following users thanked this post: Someone

Online SiliconWizard

  • Super Contributor
  • ***
  • Posts: 17799
  • Country: fr
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #6 on: February 24, 2025, 12:47:38 pm »
There are a few things to consider: Nominal already has a Sipeed Tang Nano 9K, with which he could definitely implement what he wants. It has an embedded SDRAM, which is not too difficult to implement a controller for (and there are available out there too).

Now he wanted to make his own board, so that also has constraints which may not make the Gowin FPGA route easy (as he mentioned). But in any case, I don't think that'll be wasted time. He can start by developing his ideas on the Tang Nano 9K, which is plenty for that, and once satisfied with the approach, make a custom board. And even if it's not the same FPGA, porting the code should not be hard, and SDRAM chips are available in workable packages.
 

Offline tggzzz

  • Super Contributor
  • ***
  • Posts: 23122
  • Country: gb
  • Numbers, not adjectives
    • Having fun doing more, with less
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #7 on: February 24, 2025, 01:01:01 pm »
I recommend starting with an FPGA development board; you really don't want to be simultaneously debugging hardware, plus FPGA concepts, plus PCB PDNs, plus transmission lines. Presumably you want to concentrate on your objectives.

There are many minor differences between programming FPGAs and MCUs, e.g. language syntax. I'll merely note the major differences you will need to understand:
  • everything should be parallel; think of having an infinite number of hardware threads in an MCU :)
  • clock domains, when metastability might rear its head, and how to tame its consequences
  • FSMs, and how to code them. There are many design patterns, but the "two process" pattern offers simplicity and maintainability. Process1= synchronous clocked state register, process2=asynchronous logic defining the next state
  • the building blocks primitives supplied by the FPGA manufacturer. Learn what they are, plus how to slot them into your design
« Last Edit: February 24, 2025, 01:29:25 pm by tggzzz »
There are lies, damned lies, statistics - and ADC/DAC specs.
Glider pilot's aphorism: "there is no substitute for span". Retort: "There is a substitute: skill+imagination. But you can buy span".
Having fun doing more, with less
 
The following users thanked this post: Smokey

Offline pcprogrammer

  • Super Contributor
  • ***
  • Posts: 6110
  • Country: nl
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #8 on: February 24, 2025, 01:02:29 pm »
... Sipeed Tang Nano 9K, with which he could definitely implement what he wants. It has an embedded SDRAM, ...

I find these Gowin FPGA's confusing when it comes to which memory is used. Member avst played with a Tang Nano 9K for the video capture thing we are working on and he told me that is has PSRAM instead of the SDRAM used in the FPGA on the Tang Nano 20K.

But on the Sipeed site they mention SDR SDRAM 64M bits, although the picture of the board shows a GW1NR-LV9 QN88PC6/I5, and for that one Mouser states it has abundant PSRAM.  :-//

So which is it.

Online SiliconWizard

  • Super Contributor
  • ***
  • Posts: 17799
  • Country: fr
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #9 on: February 24, 2025, 01:12:55 pm »
... Sipeed Tang Nano 9K, with which he could definitely implement what he wants. It has an embedded SDRAM, ...

I find these Gowin FPGA's confusing when it comes to which memory is used. Member avst played with a Tang Nano 9K for the video capture thing we are working on and he told me that is has PSRAM instead of the SDRAM used in the FPGA on the Tang Nano 20K.

But on the Sipeed site they mention SDR SDRAM 64M bits, although the picture of the board shows a GW1NR-LV9 QN88PC6/I5, and for that one Mouser states it has abundant PSRAM.  :-//

So which is it.

Ah, that's because it can have either, depending on the package: https://www.gowinsemi.com/en/product/detail/49/

The GW1NR-LV9 in QN88P package appears to have PSRAM indeed, with a 16-bit data bus. And from the Tang Nano 9K schematic, is its a QN88P package. So, that would be PSRAM.
Making the Sipeed Wiki bogus, which states SDRAM.
« Last Edit: February 24, 2025, 01:15:59 pm by SiliconWizard »
 

Offline Nominal AnimalTopic starter

  • Super Contributor
  • ***
  • Posts: 8349
  • Country: fi
    • My home page and email address
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #10 on: February 24, 2025, 01:56:38 pm »
If you want to learn FPGAs, than buy an evaluation board
This is not that.  I already have those for that purpose.  This is a specific hardware project with specific issues – FPGA exposing RAM to an MCU, while simultaneously itself accessing that RAM; or using two separate RAM ICs and having the FPGA use one while the MCU uses the other –, and I want to better understand those issues before starting the actual build.  I see many OSHW FPGA projects struggling with accessing external memories (although those tend to be higher bandwidth DDR dynamic memory chips).

I could pick an existing FPGA dev board with sufficient RAM on it, but that isn't designed to interface that same RAM also to a MCU.  This is a design-level issue; a typical devkit will have "sufficient" interface to RAM, and that may not be enough for me here, especially if I end up using two separate memory chips with each having their own data bus to the FPGA.

There are many minor differences between programming FPGAs and MCUs
I'm fully aware.  I've used several other description languages before, I'm not stuck in the imperative worldview at all.

(Not all programming languages describe the steps a "processing unit" should do.  There are several completely different approaches, even though most software developers only learn the imperative approach (and occasionally a functional one, like Ocaml or Haskell).  In particular, description languages describe the rules of operation, evaluating everything in parallel, instead of as a sequence of steps.  OpenSCAD implements a simple descriptive language for its constructive solid geometry: there is no progression through the code, it's all evaluated at once, all statements in parallel.  For animation, the entire "program" (geometry ruleset) is re-evaluated, once per animation frame, with a built-in variable representing the time/frame.)

I recommend starting with an FPGA development board
Again, I have those.

The question I'd like to focus on is the design and implementation aspects when interposing an FPGA between the MCU and the external RAM.
All the background and context was there because I thought it was necessary to describe why "don't do that, <do something else> instead" won't work for this use case.

The MCU access will not be linear, they will be random.  The FPGA/display part will be quite linear.  The issue is that those do not mix well, so as a workaround, I was thinking of using two separate external RAM chips, each with their own bus and I/O pins, so that while the FPGA accesses one, the MCU can access the other, with no contention.

Another way to look at this would be to consider the case where you have two (identical) MCUs, and would like them to access a shared asynchronous SRAM (so no bursts), but dual-port SRAM is too expensive so you use an FPGA as the RAM controller.

I see now that I should have omitted all that background context, and focus on the FPGA aspect, the technical details involved in using an FPGA between RAM and a MCU.  Apologies; me fail, initial post poor quality.  I will try to do better in the future.  In the mean time, although I don't like changing the "past", I think I'll add a clarifying paragraph to the question, so others will more easily understand what I'm asking about.
 

Offline tggzzz

  • Super Contributor
  • ***
  • Posts: 23122
  • Country: gb
  • Numbers, not adjectives
    • Having fun doing more, with less
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #11 on: February 24, 2025, 02:11:56 pm »
The question I'd like to focus on is the design and implementation aspects when interposing an FPGA between the MCU and the external RAM.
All the background and context was there because I thought it was necessary to describe why "don't do that, <do something else> instead" won't work for this use case.

The MCU access will not be linear, they will be random.  The FPGA/display part will be quite linear.  The issue is that those do not mix well, so as a workaround, I was thinking of using two separate external RAM chips, each with their own bus and I/O pins, so that while the FPGA accesses one, the MCU can access the other, with no contention.

What primitives are in the FPGA hardware and/or library? Dual port RAM is not uncommon.

What's the cycle time, latency, jitter and transfers/s for each of the RAM, MCU and display? Hardware FIFOs may be able to reconcile those, and are a typical library component.

Understanding the specific FPGA's hardware building blocks and manufacturer defined library components is more important than understanding the processor instruction set and software library.
There are lies, damned lies, statistics - and ADC/DAC specs.
Glider pilot's aphorism: "there is no substitute for span". Retort: "There is a substitute: skill+imagination. But you can buy span".
Having fun doing more, with less
 

Offline Nominal AnimalTopic starter

  • Super Contributor
  • ***
  • Posts: 8349
  • Country: fi
    • My home page and email address
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #12 on: February 24, 2025, 02:48:50 pm »
Understanding the specific FPGA's hardware building blocks and manufacturer defined library components is more important than understanding the processor instruction set and software library.
Right.  I haven't chosen any specific FPGA yet.  The arithmetic-logic stuff I need is on low-bit-count integers, say 6-bit values, with the major one being a multiplication-addition between two 6-bit values and obtaining the high 6 bits of the results.  To choose an FPGA well suited for being between an MCU and external memory, I'd need to understand the design aspects first, and indeed what manufacturer defined library components would be most useful here.  I don't need a softcore in the FPGA at all, as the processing pipeline is basically a fixed display controller, just with a nonstandard nonlinear framebuffer.

From your response and reading other FPGA threads here and elsewhere, it seems that most FPGA developers seem to use basically one family of ICs, and pick the problems they want to solve based on the capabilities of that chosen FPGA.  I want to do the opposite, and hear what kind of library components are particularly useful for this topology, FPGA mainly as a memory arbitrator between an MCU and a fixed-operation core, or two MCUs.

(To stave off the inevitable suggestion, I've already looked at and have dual-core MCUs.)

What's the cycle time, latency, jitter and transfers/s for each of the RAM, MCU and display?
I'd like to target 20ns or less random access single read and write cycles for the MCU, including uninterrupted back-to-back accesses (for 50Mwords/sec random access throughput).  For everything else, I'm going to live with what I can get.  For the display stuff, the access pattern is more linear (burst and prefetch/buffering possible) but with about four separate continuous data streams, with an aggregate sustained 40-70 Mbytes/sec read rate.  Even the exact RAM size is up for discussion.

So, any specifics on what to look for in this particular case?  I definitely don't need dynamic RAM, so async SRAM, SDRAM, and PSRAM (including HyperBus) library stuff I already know to look for.  What about the library components when the FPGA itself should look like RAM?  Are those even in vendor libraries, or is it so much rarer use case that those who need such, use proprietary IP for that?

Is the routing mux at the memory IC problematic?  That is, if I use two separate memory chips and buses to those, I still want to the able to flip which accesses which.  Or is faster random-access SRAM a better answer, perhaps time-domain multiplexing the FPGA and MCU accesses?  Those with practical FPGA design experience will know; I'm hoping they'll tell.
« Last Edit: February 24, 2025, 02:53:41 pm by Nominal Animal »
 

Offline tggzzz

  • Super Contributor
  • ***
  • Posts: 23122
  • Country: gb
  • Numbers, not adjectives
    • Having fun doing more, with less
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #13 on: February 24, 2025, 03:37:09 pm »
Understanding the specific FPGA's hardware building blocks and manufacturer defined library components is more important than understanding the processor instruction set and software library.
Right.  I haven't chosen any specific FPGA yet.  The arithmetic-logic stuff I need is on low-bit-count integers, say 6-bit values, with the major one being a multiplication-addition between two 6-bit values and obtaining the high 6 bits of the results.  To choose an FPGA well suited for being between an MCU and external memory, I'd need to understand the design aspects first, and indeed what manufacturer defined library components would be most useful here.  I don't need a softcore in the FPGA at all, as the processing pipeline is basically a fixed display controller, just with a nonstandard nonlinear framebuffer.

The obvious component to look for would be a memory controller! That will indicate the feasible memory types and performance.

Then just read the manufacturer's component library, understanding the type of component and how it might be used. Some components will be conceptually simple but the implementation will be tailored to the FPGA architecture - e.g. metastable-hardened synchroniser, or a FIFO which can be used to improve jitter and latency requirements/performance. Others can be complete subsystems, e.g. a DDS generator or FFT engine.

Quote
From your response and reading other FPGA threads here and elsewhere, it seems that most FPGA developers seem to use basically one family of ICs, and pick the problems they want to solve based on the capabilities of that chosen FPGA.  I want to do the opposite, and hear what kind of library components are particularly useful for this topology, FPGA mainly as a memory arbitrator between an MCU and a fixed-operation core, or two MCUs.

There's that tendency in MCU+software based designs as well.

I think you'll find that the development environment and libraries are more of a lock-in than the FPGA itself. They are very complex and there is a significant learning curve. Essentially the choice is between Xilinx, Altera, and all the niches.

Quote
So, any specifics on what to look for in this particular case?  I definitely don't need dynamic RAM, so async SRAM, SDRAM, and PSRAM (including HyperBus) library stuff I already know to look for.  What about the library components when the FPGA itself should look like RAM? 

The FPGA will be several memory addressed control/status and data registers connected directly to the MCU's IO pins; any FPGA can do that trivially.

The FPGA is also ideally suited for matching and arbitrating data transfers between MCU, memory and other components. Arbitration "clashes" will have to be accommodated by the FPGA and system architecture. There will be several different clock domains, which implies needing synchronisation components when crossing domains. Hopefully library components (e.g. FIFOs) will simplify that and aid implementing the system architecture

All that implies you need to do hardware-software co-design, just as you would with an MCUs peripherals and software libraries.

Quote
Are those even in vendor libraries, or is it so much rarer use case that those who need such, use proprietary IP for that?

Yes repeat no, depending. RTFDS!

Quote
Is the routing mux at the memory IC problematic?  That is, if I use two separate memory chips and buses to those, I still want to the able to flip which accesses which.  Or is faster random-access SRAM a better answer, perhaps time-domain multiplexing the FPGA and MCU accesses?  Those with practical FPGA design experience will know; I'm hoping they'll tell.

Providing the system level throughput/latency/jitter requirements can be satisfied, the FPGA will be able to implement that.
There are lies, damned lies, statistics - and ADC/DAC specs.
Glider pilot's aphorism: "there is no substitute for span". Retort: "There is a substitute: skill+imagination. But you can buy span".
Having fun doing more, with less
 
The following users thanked this post: Nominal Animal

Offline xvr

  • Frequent Contributor
  • **
  • Posts: 918
  • Country: ie
    • LinkedIn
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #14 on: February 24, 2025, 04:00:17 pm »
Having 2 RAMs for separate frame buffers can be a viable option if you have a lot of FPGA pins to connect 2 RAMs to and if you completely redraw each frame buffer for each frame. Otherwise, synchronizing the contents of different frame buffers will be a very painful task.
Multiple image planes can also exhaust the RAM bandwidth - each additional plane requires an additional read from the RAM for each pixel. A more preferable way to organize a plane is to use the FPGA's internal Block RAM for creating sprites (multi-level and custom geometry)
 

Offline pcprogrammer

  • Super Contributor
  • ***
  • Posts: 6110
  • Country: nl
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #15 on: February 24, 2025, 04:55:20 pm »
As an example on how to handle different streams, here is what I made so far for the video capture project. It uses block ram as semi dual port memory (one read and one write port) to interface between the SDRAM and the streams.

With burst read and write there is a lot of room in the time domain to get things done. For the streams to display this is what you need to make it work. Not an issue with static memory of course, and depending on the FPGA there might even be enough block ram to do what you want without additional memory. Depends on resolution, color depth and number of planes you need.

The SDRAM controller IP of Gowin is a bit obscure, and the manual does not provide much help, which made me decide to just write my own. Started with an example found on the net and made it fit my purpose.

SDRAM is not that hard. The three control lines RAS, CAS and WE make up 8 commands and are clocked in on the rising edge of the clock. For a read or write operation you open up a row (RAS =0, CAS = 1, WEN = 1) and then wait a couple of cycles before sending the read or write command. For a write the column address and data have to be supplied on the same clock. For a read the column has to be supplied on the same clock, but the data comes out CAS clock cycles later.

Within the same row reads or writes can be send directly, but when the row needs to be changed, you need to close the active row first with a precharge and then open up the new row and wait again before sending read or write commands.

Using burst mode the data can be retrieved or presented on a cycle basis, which makes it quick.

A refresh just has to be issued at timed intervals which is also done in my code. My setup uses about 30% of the available memory bandwidth, based on burst actions.

The SDRAM used in the GW2AR-18 devices seems to be Micron MT48LC2M32B2 64Mbit SDRAM's. They don't seem to make them anymore. https://datasheet.octopart.com/MT48LC2M32B2TG-7-IT%3AG-Micron-datasheet-8193693.pdf

For the PSRAM I can't say. Have not played with it, but it is most likely very similar from a system point of view.

Online langwadt

  • Super Contributor
  • ***
  • Posts: 5789
  • Country: dk
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #16 on: February 24, 2025, 05:05:22 pm »
for such small display why are you even considering FPGA and external ram when you can get something like this for $10
https://jlcpcb.com/partdetail/STMicroelectronics-STM32H7A3VIT6/C730222 ?
 

Offline brucehoult

  • Super Contributor
  • ***
  • Posts: 6464
  • Country: nz
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #17 on: February 24, 2025, 06:47:40 pm »
for such small display why are you even considering FPGA and external ram when you can get something like this for $10
https://jlcpcb.com/partdetail/STMicroelectronics-STM32H7A3VIT6/C730222 ?

I'm unsure exactly what that gets you for your $10.  In particular I didn't see a spec for RAM.

OP mentioned BL808, but didn't mention CV1800B (available on the Milk-V Duo for $5, which is effectively a DIP breakout board) or SG2000.

The CV1800B has 64 MB of in-package PSRAM, plus two USBs, and a 64 bit 700 MHz microcontroller. And bonus 1 GHz 64 bit Linux app processor with 128 bit vector unit.  The Duo is normally booted from a micro SD card, but under where the cars sits there are 8 pads to solder on an SPI NOR or NAND flash chip instead.

The CV1800B package is a 68 pin 0.35mm pitch QFN (7mmx7mmx0.9mm)
 

Online langwadt

  • Super Contributor
  • ***
  • Posts: 5789
  • Country: dk
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #18 on: February 24, 2025, 07:19:23 pm »
for such small display why are you even considering FPGA and external ram when you can get something like this for $10
https://jlcpcb.com/partdetail/STMicroelectronics-STM32H7A3VIT6/C730222 ?

I'm unsure exactly what that gets you for your $10.  In particular I didn't see a spec for RAM.

280MHz Cortex-M7  MCUs, 2MB Flash, 1.4MB RAM in a 100 pin LQFP
 

Offline pcprogrammer

  • Super Contributor
  • ***
  • Posts: 6110
  • Country: nl
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #19 on: February 24, 2025, 07:28:30 pm »
I'm unsure exactly what that gets you for your $10.  In particular I didn't see a spec for RAM.

On the ST site is states 1376KB of SRAM.

https://www.st.com/en/microcontrollers-microprocessors/stm32h7a3vi.html

OP mentioned BL808, but didn't mention CV1800B (available on the Milk-V Duo for $5, which is effectively a DIP breakout board) or SG2000.

The CV1800B has 64 MB of in-package PSRAM, plus two USBs, and a 64 bit 700 MHz microcontroller. And bonus 1 GHz 64 bit Linux app processor with 128 bit vector unit.  The Duo is normally booted from a micro SD card, but under where the cars sits there are 8 pads to solder on an SPI NOR or NAND flash chip instead.

The CV1800B package is a 68 pin 0.35mm pitch QFN (7mmx7mmx0.9mm)

For me QFN is a bitch to solder, but with a hot plate it might be easier, and that is what Nominal Animal has, according to his posts.

But for display interfacing I think it might be a bit limited. The chip supports mipi for video input, but not for display output. The bigger SG2002 does have mipi output support.

I think that a plus for the FPGA is the flexibility on connecting what ever display you like.

Offline asmi

  • Super Contributor
  • ***
  • Posts: 3345
  • Country: ca
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #20 on: February 24, 2025, 08:06:42 pm »

As to the display modules, I have a few 320×240 and 480×320 IPS TFTs from BuyDisplay/EastRising.  I personally also like the smaller 1"-2" ones, but for nontechnical people, the display needs to be 2.4" – 4" to be practically useful.
As I understand, all these displays already have controller onboard, so you can connect them to MCU directly via SPI-like bus, no need for FPGA for that. FPGA is typically required to drive "dumb" panels which needs to be continuously fed with data to display as they can not display anything by themselves, though higher end MCUs often have peripherals to do that too.

Offline Smokey

  • Super Contributor
  • ***
  • Posts: 3881
  • Country: us
  • Not An Expert
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #21 on: February 24, 2025, 09:13:21 pm »
As someone who has specced out the design for a product that does almost exactly what you are describing, I would recommend looking at off the shelf single part display driver solutions (memory/processor/interface/etc) first before trying to cost optimize with the least expensive separate micro, memory, and fpga hardware.  Yes, it may be more expensive per unit but unless you plan on making 10,000 of these you will never make up the difference with the extended hardware and software (and hdl) development time and effort (especially if you are starting from zero).

Before you start anything, make a spreadsheet and quantify all the numbers for what you are actually trying to do.  Display resolution, bits per pixel, buffer sizes, data transmission bandwidths between micro-memory and memory-display, display update rate, serial/parallel busses.  Don't get caught off guard by either the actual memory requirements or the data transmission bandwidths to support your desired frame rates.  Higher resolution displays need bigger pipes than you would think at first and you don't want to be either surprised or backed into a design corner.

One more thing... the display interfaces and configuration settings are almost never plug and play (especially if you get into DSI), and made exponentially worse by poor data sheets (there is a reason stuff like raspberry pi only support specific select LCDs despite having an interface with lots of options).  Expect some pain and suffering getting here too.  BuyDisplay has decent support for the space, but they only go so far technically if you only look like you are going to buy a couple displays and get too needy. 
« Last Edit: February 24, 2025, 09:16:56 pm by Smokey »
 
The following users thanked this post: Someone

Offline bson

  • Supporter
  • ****
  • Posts: 2771
  • Country: us
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #22 on: February 24, 2025, 09:15:28 pm »
I would pick a fast SRAM and implement some sort of interleaved dual port interface.  The reason is you will use up pins in a hurry with multiple memory buses.  You could use two SRAM packages and interleave them on the same bus (just mutually exclude CS's) but that doesn't really buy anything over a single package split into two logical memories.  You're already looking at maybe 25-30 pins per bus.

Make your life easier by picking a high speed grade, low complexity (low LUT count and features) FPGA.  I'd also suggest sticking with one of the big vendors for now as their integrated tools, crappy though they often are, significantly lower the learning barrier. And let's be frank, you can spend a LOT of time without any hardware, just designing modules and test benches, running simulations, and understanding how to use constraints effectively.  So you could just pick something like a MachXO2-640-6, install Lattice Diamond, and go to town - just pick a 100 pin device and pretend you're going to use it.  It's pretty easy to change devices, or have multiple devices or speed grades (another learning barrier) in a project.  A lot of QFP parts available.  The big tools have reasonably complete (though extremely scattered and normalized) help trees, application notes, and sample projects to mine for ideas.

I think the easiest is just to set up a toolset, pick a device, and start designing.  Find out what works and what doesn't, what the tradeoffs are.  Verify with test benches.  Debug using printf's and wave diagrams.  You don't actually need any hardware for this.

I have a couple of "learning" books, and they are so elementary as to be useless, so I won't name them.  They're more about introduction to logic design rather than how to implement logic in FPGA - I mention this since I think you'd done plenty of logic design before, or at least understand how it differs from software for a Von Neumann machine.  If you have logic circuit design experience you're going to love FPGAs.
 

Offline Someone

  • Super Contributor
  • ***
  • Posts: 6071
  • Country: au
    • send complaints here
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #23 on: February 24, 2025, 10:01:54 pm »
If you want to learn FPGAs, than buy an evaluation board
This is not that.  I already have those for that purpose.  This is a specific hardware project with specific issues – FPGA exposing RAM to an MCU, while simultaneously itself accessing that RAM; or using two separate RAM ICs and having the FPGA use one while the MCU uses the other –, and I want to better understand those issues before starting the actual build.
Did you know a good portion of FPGA design are prototyped on dev boards before any hardware is designed/built? While you don't have an architecture planned out (or as above calculated in terms of bandwidths) it is too early to be thinking of what parts are going into the mix.

Just like in software you can build out some stubs/simulations of bits of the system that are large/complex and test the bits you are unsure of under some limited/basic conditions, to check if it the concept is feasible.

What has been repeated and re-inforced in this thread is "two memories, exclusive access by two masters" but that does not match up with your desired single unchanging background plane with dynamic planes being updated over the top. Your description leads towards having multiple buffers for each of the planes that can be swapped in and out per plane, which isn't hard in a single memory address space.

Pointers to existing projects interposing an FPGA between a MCU and RAM (SRAM/SDRAM/PSRAM)?  Especially if some kind of access delegation is used.  This is the part that is most likely to bite me in the butt, so any pointers how to do this right would be welcome.
Having multiple [arbitrary description of separable operations] accessing a single memory is the dominant paradigm of software, and naturally extends to what you are doing. The FPGA/hardware specific counterpoint is synchronous (or close to) streaming which might only be applicable to the pixel output pipeline.

Where is the blockage in connecting [arbitrary write only] master and [arbitrary read only] master to a conventional memory? Ignore the specific interfaces they use for the moment, break it into understandable bits. If that does not make sense then adding multiple readers (and/or writers) will really get you confused.
 

Offline Nominal AnimalTopic starter

  • Super Contributor
  • ***
  • Posts: 8349
  • Country: fi
    • My home page and email address
Re: Hobbyist: FPGA between MCU and external memory?
« Reply #24 on: February 25, 2025, 05:36:35 am »
OP mentioned BL808, but didn't mention CV1800B (available on the Milk-V Duo for $5, which is effectively a DIP breakout board) or SG2000.
Yup.  I already have a SG2002 (Milk-V Duo 256M), its "big brother", too.  CV1800B would do absolutely fine for me, but I don't see any hobbyist designs to compare to.  The schematics (PDF) and docs are available, but I'm not sure if I can pull that off as a hardware design.  I fear I'd be bitten by the power-up sequencing of the four supply voltages it needs (3.3V, 1.8V, 1.35V, and 0.9V).

The CV1800B package is a 68 pin 0.35mm pitch QFN (7mmx7mmx0.9mm)
I suspect my soldering results for 0.35mm pitch QFN would be even worse than they are for 0.8mm pitch BGA. :-[

(Even the early versions of SparkFun MicroMod Teensy, using an M.2 connector for peripherals, suffers from contact issues due to the PCB bending a little.  Using a thicker PCB would help, but standard M.2 connectors limit the thickness.  I don't like the idea of trying to do single boards better than SparkFun can design and manufacture in batches.)

As I understand, all these displays already have controller onboard, so you can connect them to MCU directly via SPI-like bus, no need for FPGA for that.
Yes.  The ones I like most use ILI9341, ST7789, ILI9488, and similar controllers (they have only minor differences between them) exposed via 40-50 pin FPC.

In the overlong background, I explained that I already have flat framebuffers in hardware, but would prefer a layered one instead, with one full-color background framebuffer, with two or three indexed-color/paletted overlay framebuffers with full alpha support on top.  I can do this in software on a sufficiently high-powered ARM microcontroller, but the composition is a parallelizable operation best suited for FPGAs and ASICs, not general-purpose processors.  I could do this the retro way, using discrete logic and fake ALU chips (MCUs or SRAMs or EEPROMs), all running off a 12-18 MHz bus clock, combining data from 3-4 parallel data buses (retro way using SRAM chips with address coming off a counter).

your desired single unchanging background plane with dynamic planes being updated over the top
No, they are all mutable.  I could show the software implementation of the approach, and how much this reduces memory use and required memory bandwidth if the composition of the planes is done by external hardware.  Simply put, in my use cases and omitting the composition cost, the memory use and number of operations needed per displayed frame is a small fraction of that needed when using a standard flat full-color framebuffer.

Even more so, if the plane is slightly larger than the visible area, and the start address is programmable within a fixed-size wraparound region, for implementing scrolling.  Only the new data becoming visible will need to be written, old visible data does not need to be touched.

It is a waste of resources to recompute each pixel every frame if you only have one general-purpose core, though.

I would recommend looking at off the shelf single part display driver solutions (memory/processor/interface/etc) first
Like I wrote in my overlong initial post, I have already implemented this on a microcontroller, and have played extensively in software implementations of such nonlinear framebuffers.  It is not a question of whether I can do it, but whether I can do it "better" than what I already have, within my oddish constraints.  My budget is close to zero because of personal situation, but I have no time limit at all.  I've played with this for years now on ARM MCUs as a hobby already.

It can sound odd that I want a "real" FPGA project when I haven't properly learned any HDL yet (target being Verilog-2005 due to Yosys support).  Thing is, I know how I learn best, and having an example use case to apply my understanding to is very useful.  Choosing the complexity level is tricky, because I quickly lose interest when the problems are easy to solve; but can chew on them if they are difficult but not clearly intractable for a single hobbyist.  Don't worry, I will not be limited to the subdomain that covers my immediate needs, though: that only acts as my opening, my stepping stone, into understanding the entire approach, the worldview, the paradigm, with my curiosity driving me to expand outside my immediate needs.  This just works very, very well for me, and helps me keep an open mind, not letting my existing conceptions hinder me.  Not everyone learns like this, though, so this approach may seem ass-backwards or even doomed to failure to others, just because they learn best some other way.

Just like in software you can build out some stubs/simulations of bits of the system that are large/complex and test the bits you are unsure of under some limited/basic conditions, to check if it the concept is feasible.
Of course.  I've got a "library" (mostly in offline storage) of thousands of such unit test cases in software, examining specific algorithms and math, accumulated throughout the years.

I do believe I'll start by connecting one of the FPGA dev boards I have to one of my Teensy 4.1 MCUS, with the FPGA pretending to be PSRAM.  Teensy works very well with APS6404L_3SQR QSPI PSRAM at 84 MHz clock, and I can lower the QSPI clock to compensate for leads-in-the-air wiring (but I'll try to use the every-other-wire-is-GND approach too).  For read tests, I can use the address bits reordered and XOR'd to generate the data, and verify all sorts of read access patterns by the MCU yields correct results.

Next step, I can implement read-write access by using internal RAM (for example, using only a few least significant bits of the address).

Separately, I can connect an APS6404L PSRAM or more, each with their own QSPI bus to an FPGA, and that FPGA to a Teensy 4.1, and test the case where the FPGA is a simple multiplexer.  For control, I think I'll use an UART, with one-byte commands from the MCU to the FPGA, with the FPGA responding asynchronously with one-byte state information when the command is completed.  I need the asynch'ity later.  I'll also do the UART FPGA implementation separately first, before combining with the PSRAM muxing.

This is how I do software development in general, too: I implement subsystems in isolation, but with a view of future integration/adaptation.  After verifying the subsystem works in isolation (often using a test harness), I integrate the subsystem to the separate combined project, followed by rigorous testing.  (I do not usually create software libraries per se, as I end up adapting each previously tested approach for a new use case.  I do not copy-paste or just include my old libraries; I've found I get better results when I adapt each tested approach to the needs of each actual use case.  Testing each adaptation thoroughly before integration is very important, though: bugs will occur, but limiting them to a small changeset avoids a lot of frustration and saves time.)

I'll do the FPGA-only display controller separately, too.  I already have a Tang Nano 1K and Tang Nano 9K with compatible display modules to do this with.  I've also got the schematics for these, and Mouser sells the FPGA ICs, and the basic schematic seems simple enough that I should be able to design my own board with these later on.

As a next step to that, I can check if APS6404L_3SQR PSRAM (in linear bursts) as a data source for the framebuffer is fast enough, and get a ballbark idea of the kind of resources and data bandwidth needed from the FPGA for this aspect of the future whole.

A sensible plan seems to be building up nicely, thanks everyone!  :-+
« Last Edit: February 25, 2025, 05:41:32 am by Nominal Animal »
 


Share me

Digg  Facebook  SlashDot  Delicious  Technorati  Twitter  Google  Yahoo
Smf

 

-->