The question is the difficulty level for a hobbyist to use cheap FPGAs to expose SRAM or SDRAM or PSRAM to a microcontroller. This would be my first real FPGA project, and I'd like to avoid BGA, too. If you don't already know, I'm a lightweight very-limited-budget hobbyist in electronics, total newbie with FPGAs, but have decades of programming experience (including some low-level, bare metal stuff) in about a dozen programming languages.
I have a dual purpose for this thread: one is to perhaps get advice on starting this hobby project, and the other (more important to me) is to discuss similar use cases and projects others might be working on. I don't know how to simplify this starting post without losing context, so sincere apologies for those who hate the walls of text!
Edited to clarify:
The core of the question is about the technical and especially design aspects of putting an FPGA between a memory IC and a microcontroller. Or, equivalently, an FPGA with built-in memory between two identical microcontrollers; the MCU or MCUs see only ordinary external memory. In my case, I don't necessarily need access arbitration, because I can use two RAM ICs with separate buses, with each MCU accessing their own one, never the same RAM, as long as I can swap which one they access. I believe this is a relatively unusual design, with the FPGA-uses-own-RAM being the common one, with only rarer devices like oscilloscopes and other data acquisition equipment sharing memory this way, so I want to learn about the important design-affecting issues and approaches up front.
The rest of this post is details on what and why, but the above is the part that gives me pause.
BackgroundI like to interface small 1" - 4" display modules to Linux-based appliances and single-board computers for use as an external information display for nontechnical humans. High-fidelity IPS panels up to 800×480 (my preference is 320×240) are quite cheap. I can do this using various microcontrollers with USB connection to the host, with the host software only controlling the display, not generating the image. The issue is that typical microcontrollers only expose a single framebuffer, and that's just not very effective. The displays have their own controllers and easy digital interfaces, so I find the idea of using an FPGA as a programmable framebuffer extremely interesting. I'm only a hobbyist, so I want to keep the costs down and everything hand-solderable, so the optimum would be to use external SRAM/SDRAM/PSRAM as a framebuffer exposed to a microcontroller to draw onto, with the FPGA updating the display module transparently on the background.
In particular, I like the idea of having 3-4 completely separate planes: a background one with full 16/18-bit color, plus 2-3 various indexed/paletted planes drawn on top. Similar stuff was used in arcade games and home game consoles, back when memory was scarcer. (Even today, with our stupidly powerful GPUs even in phones, the actual displayed image is composited from separate "sub-frames" in graphics memory.)
A practical example of this is a small gauge display. The static part is in the background plane. The gauge needles and displayed values are on a separate plane on top, so the MCU does not need to modify –– read-then-write –– to change these, just clear and draw the new ones on top. Because the memory for each plane is separate, clearing the gauge plane does not affect the full-color background at all. Plus, the MCU does much less work.
I often use capacitive touch to control what is displayed, so yet another plane can be used for the touch interaction; I like rings or circular arcs extending outwards from the touch point like a ripple to indicate touch registration (separate from touch action on the graphics). As a separate plane/framebuffer in memory, the ripples can just be copied to the framebuffer by the MCU without any calculation, or any effect on what is being done to the other planes.
If there is a fourth plane available, it can be used in between, for additional overlays and transients. The indexed color mode I'm most interested in these is one byte per pixel, with each of the 256 possible values freely chosen from a full-color palette (262k colors) with a full transparency control. For creating antialiased text and polygons, using 3 or 4 low bits for transparency makes things easier for the MCU, but precise and clean graphics and antialiased text.
It would also be fun to play with the kind of visual models used in old game consoles. All my hobby stuff like this is CC0-1.0/Public Domain, so if I get this working, it'd also be available for other hobbyists for use in old-style portable pixel art and retro gaming projects, too.
As an end result, I'm hoping to end up with a display carrier board with an FPC connector to the display modules using ILI9341/9488/ST7789 etc. controllers, containing the MCU, external RAM, FPGA, USB connector, and perhaps a microSD socket for graphics asset storage. I currently only use Linux and open source tools, so I'd prefer to avoid proprietary software; and I don't have funds for developer kits.
Hardware:Among the many microcontrollers from NXP, STmicroelectronics, and others, I've taken a liking to
WCH CH32V307VCT6: it has a high-speed USB 2.0 (480 Mbit/s) and an external parallel static memory interface (FSMC) interface for up to 16 MBytes (but with data and address bits multiplexed on the same pins). It's cheap (<5€ in singles), programmable using open source tools in Linux, simple enough for a hobbyist like myself to put on a custom board, and although not terribly powerful, definitely powerful enough for this.
As to the memory itself, I've been thinking about using two SRAM modules, something like IS61/64WV51216EEBLL 512k×16 SRAM (
PDF datasheet at Mouser), with one "connected" to the MCU while the FPGA reads the other for the framebuffer generation. As I've not done any real projects with FPGAs yet, this is purely to keep the two separated, so that I don't have to work out a bulletproof sharing logic between the FPGA framebuffer read access and MCU read/write access. Having access only to the part of the framebuffers not currently being displayed is absolutely fine for me, even if it means extra cost: this is hobbyist one-offs, not production stuff. I do believe CH32V307VCT6 memory interface has a WAIT input I might use to interleave FPGA and MCU accesses, but I know how risky that kind of complexity is in a hobby project, so I'm aiming to keep it simple, even if it means extra hardware. The downside is that using two in parallel would take quite a lot of pins: about 40 per SRAM module, another 24 to the MCU, and another 24 or so for the display module.
Even faster SDRAM might make things easier. I do need very fast random access for the MCU (20 ns write cycle time or less). While the FPGA needs similar throughput (I'd like to get about 70 Mbytes/s), it's almost all just linear scanning. With fast enough SRAM, the FPGA could probably use fixed time division multiplexing approach, using "access slots" for the MCU, the FPGA frame buffer, and SDRAM refresh. I know nothing about how to control SDRAM or any of the other RAM types; SRAM thus far seems the simplest to implement. I'm very open to suggestions, though.
As to the FPGA, I have no idea. I prefer to avoid BGA, which limits my selection, especially as I'd like to have > 100 I/O pins. I'd like to have fast 6×6=12-bit multiplication, because optimally I'd do 24 of those in parallel in the output stage (per pixel), with a throughput of about 10 Mpixels/second. Otherwise the logic is limited to marshalling the RAM accesses and RAM lookup. Many FPGAs seem to support different voltage levels on different banks. The MCU and display I/O will be 2.8V–3.3V. Again, this is my first real FPGA project, so any suggestions would be appreciated.
As to the display modules, I have a few
320×240 and 480×320 IPS TFTs from BuyDisplay/EastRising. I personally also like the smaller 1"-2" ones, but for nontechnical people, the display needs to be 2.4" – 4" to be practically useful.
Existing software implementation:I already have experimentally implemented the image generation stuff on a computer and on an MCU, so it's just the hardware that gives me pause. I have considered using a dual-core microcontroller, with one core dedicated for the display generation and updates, but most of the cheap ones just do not have enough RAM for multiple framebuffers/planes, or they support full-speed USB 2.0 only (12 Mbit/s, not high-speed 480 Mbit/s) –– for example, Raspberry Pi RP2350 and Bouffalo Labs BL808 ––, and I feel those are limits I don't want.
(I also have a largeish collection of pixel tricks and effects dating back to nineties. I particularly like retro arcade and puzzle games with orthographic projection views, like in Gauntlet and similar, and have some tricks for those if anyone is interested.)
The software is all experimental code, not published yet, but I like this "new" approach (well, going back to roots of computer graphics, retro game consoles and the like, except with better hardware) A LOT. So, this is not a question of "is this possible", but "how could I do this in hardware", and have fun at the same time. And hobbyist-level stuff only! I don't care if someone else develops this further into a product, but I definitely will not. If the end result works, I'll put it all out in OSHW, likely CC0-1.0/Public Domain.
(CH32V307VCT6 has only 64kBytes of RAM, which does not suffice even for a small full-color framebuffer. The FSMC interface is not pin-compatible with any SRAM chips I've found, because of the way it multiplexes the address and data pins. I could maybe design a circuit for the needed glue logic, or just use the host computer/appliance and just stream the framebuffer changes.. but I don't want to.)
My favourite microcontroller,
Teensy 4.1 (based on NXP i.MX RT1062), has a high-speed USB 2.0 (that even with just plain old USB serial in Linux can sustain 25 Mbytes/s in one direction), lots of RAM (1 MBytes) plus PSRAM support (up to 16 MBytes), and has a Cortex-M7 running at 600 MHz so it has no issues doing this in software. It's pinout isn't optimal for driving these external display modules, and although PJRC sells the proprietary bootloader chip (which also handles power sequencing on startup) so I could design my own board with a more betterer pinout, it is 0.8mm pitch BGA, and I still fail at BGAs. In other words, I can do this on relatively cheap hardware already; it's just that an FPGA-based reprogrammable hardware approach exposing the framebuffer/graphics resources as external memory would be even more useful.
Assuming it's not expensive, of course.
Advice needed:Suggestions for the FPGA? I know of Lattice iCE40, and have a Sipeed Tang Nano 1k and 9k development boards. So, a complete beginner. I'd like to spin my own boards, and avoid BGA. The displays are in the 2.4"-4" range, so board space is not an issue; I'll use JLCPCB or similar for the boards. I'm not interested in expensive developer boards and kits, because I have a tiny budget.
Pointers to existing projects interposing an FPGA between a MCU and RAM (SRAM/SDRAM/PSRAM)? Especially if some kind of access delegation is used. This is the part that is most likely to bite me in the butt, so any pointers how to do this right would be welcome. I only need a couple of megabytes of RAM total, so no need to go into dynamic RAM and their intricacies: I'd like to keep this simple.
Advice, articles, projects, or books showing how to interface SRAM/SDRAM to an FPGA correctly? I like async SRAM because it is simple. SDRAM requiring periodic refresh cycles scares me a bit, especially because the MCU will not be aware of those. I can control the async timing on the MCU, but I'd rather keep it as fast as possible to make maximum use. (The CH32V307 maximum clock frequency is 144 MHz, but I'd be happy with 20ns read and write cycles to the managed external memory.)
Note the tiny budget, though, and this being hobbyist only, and that I'm a Linux-only user right now: I do not have any Windows licenses, and relying on Wine feels clunky, so any Windows-only toolchains are not possible for me. I'm hoping I can find something supported by Yosys, nextpnr, and similar open source projects, although I'm happy with free vendor tools too, as long as they run in Linux.