6) If you want to test a HDL implementation, you can simulate the actual portable synthesizable HDL either on a PC (verilator, icarus verilog, vivado, quartus, ...) and if the algorithm is as simple as you say it'll probably run the HDL simulation fast enough to be real time on a decent PC or at least fast enough to prove the concept with good test vector full frame clips data sets.
Hello everyone!
This will be my first project using an FPGA so please bear with me. I'm trying to implement an artifical retina on a Spartan 7 FGPA (Cmod S7). I know this won't cut it for the end system and that I need to step up to a Zynq 7000 in the end, but I want a cheap proof of concept before spending too much money.
I intially implemented the retina in software (using C and Python) which revealed the major bottleneck to be the memory speed! The algorithm itself is extremely simple (just multiply accumulate and addition) which as far as I know can be done in 1 cycle uaing the DSP slices in the FPGA. My idea to overcome the memory bottleneck was to use many QSPI NOR flash modules to create a wide bus (64 or maybe 128 bit wide bus). The only operation on these NOR flash modules would be read operations which should be fast enough for this apllication (I'm aware they are very slow to write to, so they will only be used as a ROM).
The accelerator needs to recieve an image as an input and return the result to the program. I'll be using the USB port for proof of concept which is oit ideal and I think the AXI interface could be used if I go with the Zynq solution.
Please let me know if this idea is feasible or if I'm thinking about it the wrong way around. I will add some more details about the exact details of the algorithm, the software implementation and how I plan to modify it to work with an FPGA in the next post (I have to simplify things so they can fit here).
P.S: I forgot to mention that the total memory usage is around 300MB, so I can't rely on BRAMs. DMA with DDR memory would be next best option, however that won't address the bandwidth issue, which is why I want to use sperate modules to create a wide bus.
Granted I didn't mean real time in the sense that it'd be as fast as peak FPGA hardware for decent FPGAs.
@cython.wraparound(False)
@cython.boundscheck(False)
cpdef sample(unsigned char[::1] img_flat, unsigned short[::1] coeffs, unsigned int[::1] idx, unsigned int[::1] result_flat):
cdef unsigned int x
with nogil:
for x in range(img_flat.shape[0]):
if coeffs [ x ] > 0 :
result_flat [ idx [ x ] ] += img_flat [ x ] * coeffs [ x ]
The big 200W eating beasts of a GPU used to run games and mine Etherium coins have incredible amounts of processing power when compared to any CPU. The FPGA implementation of a computational machine with similar processing power would likely require a large number of 10k$ FPGAs on a board while burning >1kW of power to run. You most likely do not need this level of power.
....
FPGAs might seam low power because they consume so little power (Most of the time they don't even need heatsinks) but they have only a small fraction of the computational performance of a GPU. Getting high GB/s of external memory bandwith out of a FPGA is also not trivial.
I intially implemented the retina in software (using C and Python) which revealed the major bottleneck to be the memory speed!
1) Use something like a KRIA dev board which is specifically made for easy prototyping of image / video processing?
https://www.xilinx.com/products/som/kria.html
The memory IS CONFIRMED to be the bottle necking factor (I have received hundreds of replies regarding this). Also FPGA offers advantages like ability to prefetch data into buffers, whereas a CPU has to buffer a small amount of data into cache, process that data, send a request to memory to fetch more data and sit there idle for many cycles (which is made even worse due to memory latency. Some of this can be addressed with things like out of order execution, branch prediction, etc. but despite the on paper specs, FPGAs do very well, sometimes even better than a desktop PC for image processing tasks. Here is a quick demo of Zynq FPGA beating a computer in a very similar task:
Please read the previous replies before posting an answer.
.. The memory IS CONFIRMED to be the bottle necking factor ...
You might also consider SRAMs since they're fast and wide and deep enough that they might be useful for your cache level processing if such won't fit in BRAM.

The memory IS CONFIRMED to be the bottle necking factor (I have received hundreds of replies regarding this).
I have attached my dissertation where you can find the full details of the algorithm itself (chapter 4 is the design section and is relatively short, about 3 pages). From the results section you can see that multi threading adds no performance benefits but increasing the memory bandwidth (and even reducing its latency) scales linearly.
The memory IS CONFIRMED to be the bottle necking factor (I have received hundreds of replies regarding this).Then justify this, a specific implementation on a CPU there with a trend connected to memory performance is not conclusive. You may have wandered around a local minima and missed the bigger picture. Given the sparse nature of the convolution output and the symmetry/scaling of the kernels, you've probably chosen a poor implementation method. Also, what works well on a CPU is not what works well on a FPGA, so an entirely different approach may be needed.
What are the fundamental parameters that are causing memory pressure? Rates for pixel input, kernel lookup, pixel output, their associated bit depths, these can be quantified.I have attached my dissertation where you can find the full details of the algorithm itself (chapter 4 is the design section and is relatively short, about 3 pages). From the results section you can see that multi threading adds no performance benefits but increasing the memory bandwidth (and even reducing its latency) scales linearly.Just to reiterate the "application" so we're talking about the same thing: Its a sparse convolution to make an irregularly sampled image from a regularly sampled one?
Usually thats only done for simulation of what an optical system could provide as the raw data to a further processing system (such as a neural network), doing it in realtime is a layer of abstraction that is almost certainly wasteful as the downstream processing system could be trained on a "cheaper" representation of the data and end up with a similar result. Or, given a power/processing/cost budget, putting more resources into the inference/intelligence and less into making it "nice" from a bio-motivated or conceptual ideal. Ditch the perfectly sized/shaped gaussian type kernels, what happens when the inference is trained/tuned on simple decimation from rectangular averages?
Spending huge time and resources to change the data format from uniform to this arbitrary and complex format without considering other ways to achieve the desired result of the system is a big waste of time. I know, I've worked on this exact field, and resolved this exact problem of resampling regular images for downstream processing. You may have been tasked with this little slice of a larger project by someone else, or misunderstand the significance of foveated image structures, but its pushing to build a very complicated system that doesn't solve any real problem.
The cost of doing an ideal foveated data compression/reduction step is disproportionate to its possible improvements elsewhere.
...
For such a simple 24-pin IC with 12-data lines or such it wouldn't be hard to just make a plug in module or PCB yourself as needed holding from 1-4 devices.
There's one example of a RAM adapter board here:
https://docs.icebreaker-fpga.org/hardware/pmod/hyperram/
1) Use something like a KRIA dev board which is specifically made for easy prototyping of image / video processing?
https://www.xilinx.com/products/som/kria.html
I guess for $200 it would be best bet for a cheapest solution.
and while awaiting a kit, familiarize how to
https://github.com/Xilinx/Xilinx_Kria_KV260_Workshop
https://www.xilinx.com/support/documentation/white_papers/wp529-som-benchmarks.pdf
.. The memory IS CONFIRMED to be the bottle necking factor ...What throughput and latency are required for your application?
this process is still the main bottleneck of a CNN which uses fovea sampled inputs
The idea of just averaging rectangular areas might also be worth investigating but outside the scope of the current thread. I'll try to look into it, but I actually feel like it's going to perform worse (changing MAC with accumulate and divide, which as far as I know is not an instruction. Also division is about 3x slower than multiplication (at least on mainstream intel platforms)).
Regarding the data throughput problem: I have confirmed it via 4 separate methods, only one of which (comparison between two different speeds, or as you call it the trend) was mentioned in the paper. Other methods were using theoretical maximum speed calculations, memory throughput benchmarks and finally an alternative packing method which processes the exact same data, but gets rid of blank elements to reduce memory throughput, which scaled linearly. The results are very conclusive if you ask me.
A better approach would be creating a larger image that contains multiple kernels, which can be thought of as a larger kernel; This approach only needs one calculation to find the corner coordinates and allows for a cache friendly way to multiply the original image with this new coefficient image.
Thats the problem, creating synthetic foveated data with this method is expensive and without any justification for its accuracy/quality. The CNN likely doesn't care and can be retrained on cheaper data sources. Trying to make synthetic foveation faster/cheaper is almost certainly an inefficient use of time/effort. Give the CNN an equally arbitrary layer there and it will find a better solution, thats the whole movement to ML and CNNs. Computers find better convolutional structures than human designed/imagined ones, put as much inside the ML as possible.
Assuming the foveation is entirely done with radially symmetric gaussians (or at least small shapes that can be repeated with reflection) then your primary claim is incorrect:QuoteA better approach would be creating a larger image that contains multiple kernels, which can be thought of as a larger kernel; This approach only needs one calculation to find the corner coordinates and allows for a cache friendly way to multiply the original image with this new coefficient image.That might be simpler from a high level language and may well be more efficient if the abstraction overheads are quite high in your test. But in FPGAs that no longer applies and there are huge gains from doing things with smaller localities of memory/data. Your tests are only valid for the assumptions/platforms/structure that you haven't fully defined, and are therefore likely to be misleading or entirely incorrect when applied to a different computational architecture. "one calculation" is nonsense when compute resources are counted in add/mult operations, memory bandwidth is counted in bits/s, but you won't tell us what the actual compute task is.
A gaussian 5x5 is not the naive 25x mults + 24x two way adds + 50x 32bit memory lookups. It might be in a very poor implementation, but even a trivial FPGA architecture is 6x mults (some can be simplified out with binary shifts) 24x two way adds and streaming data (zero memory lookups). This can be generalised to support arbitrary sizes and sub pixel positioning, still at lower computational cost than brute force convolution. This is all routine work that has been done many times before by many different people.
Thats the problem, creating synthetic foveated data with this method is expensive and without any justification for its accuracy/quality. The CNN likely doesn't care and can be retrained on cheaper data sources. Trying to make synthetic foveation faster/cheaper is almost certainly an inefficient use of time/effort. Give the CNN an equally arbitrary layer there and it will find a better solution, thats the whole movement to ML and CNNs. Computers find better convolutional structures than human designed/imagined ones, put as much inside the ML as possible.I feel like you're entirely missing the point of "foveated vision". To process large images (1080P for example), a CNN will be nowhere near fast enough due to the massive amount of data. A retina samples this down to a much lower number of elements (for example 50K), which is smaller than even a 640x480P image (around 307k). You may spend more time on the foveation step compared to a simple convolutional layer, but you massively decrease the amount of data that needs to be processed. It's essentially a lossy compression algorithm.
Cropping or resizing have their own downsides which are mentioned in the dissertation.
The methods I used do seem weird or even inefficient, but have been the result of about 1.5 years of experimenting with different approaches and are proven to be the fastest (CPU rendering that beats GPU performance, just think about that).
The wrong way of approaching any challenge is: "Well I might as well not try since I have no clue what I'll be doing"; that way you'll be stuck forever where you are. Even the most skilled programmers or god like engineers, started from not knowing anything and learned their way out of the problems they faced. I'm certainly up for the challenge anyways.
I have clarified this several times on this thread but people seem to want to skip over it or may not understand the end goal. I'm not trying to detect peopleThe memory IS CONFIRMED to be the bottle necking factor (I have received hundreds of replies regarding this).
Even if the end product is slower than a top of the line desktop PC, it's fine since the target to beat is a raspberry pi, not an AMD 5950xPlease read the previous replies before posting an answer.