EEVblog® Electronics Community Forum
Electronics => FPGA => Topic started by: ebastler on June 13, 2020, 09:04:49 am
-
I have a nearly completed VHDL design for the Spartan 6, implemented via ISE 14.7. The automated synthesis and place/route process reliably produces a design which runs at above 90 MHz, but then things get erratic. (Yep, 90 MHz are not a whole lot -- the major bottleneck is that I am using all block RAM as a large 64 kByte RAM, so there are long routing delays to and from the outer blocks.)
I would like to push this to 100 MHz, mainly because that's a three-digit number. ;)
I can come close, and did get timing closure at 100 MHz once, but the optimization results have been frustrating. It seems pretty clear that the place & route does not always find the optimal result. E.g.:
- I set the internal clock to 94 MHz, and the timing results tell me that my design can run at up to 98.x MHz. So I set the clock to 98 MHz -- and get a design that only runs at up to 92 MHz. I did check the clock jitter and it did not increase with the adjustment to 98 MHz, so that's not the cause.
- Trivial changes to some VHDL signals, which clearly have nothing to do with the critical paths, also can throw the result off track significantly.
- I switch the clock generation from PLL (which is power-hungry) to DCM_CLKGEN, which makes the clock jitter a mere 25ps worse. But the penalty on the resulting overall design clock speed is nearly 10 MHz?! Both clock generators are directly driving a BUFG, so I would assume clock distribution behaves the same for both of them.
Timing and placement constraints are probably the answer to my frustration, but I struggle with them. Based on earlier advice here, I have declared some false paths for signals derived from stationary inputs, and at least that seems to do no harm --- although I am not sure about reliable improvements. I have also tried placement constraints to put some critical functionality fed by the BRAM close to the chip's center, to ensure it's equidistant to the outermost BRAMs on both ends of the chip. This helps in one particular place/route run, but then when I change some other trivial detail, the placement restriction actually makes the results worse compared to a run where I remove it.
What's the right way of doing this? I have not found any good tutorials with specific advice on how to use constraints in a robust way. Thanks for any hints you might have!
-
What you describe is what happens when you are on the edge of what performance you can expect from the device, given your design. When you turn up the clock to 98MHz, the tool will give up faster when it figures out it cannot be done, or it has already tried different strategies that turned out to be worse.
Your frustrations with this will not end unless you make changes in the code itself. It is often a question of low latency vs. high bandwidth - and it seems that you have already chosen low latency for fast access to your BRAM. And it will only get worse when you start filling up the device with more features.
You could a) pay more money for a higher speed grade device, or b) Tweak your design towards higher bandwidth, by adding more registers/pipelining in your critical paths. Also consider utilizing full-bandwidth elastic buffers (double-registers or SRL16's) to build tiny FIFOs which consumes very litt resources.
-
You can read this if you haven't already: http://www.xilinx.com/cgi-bin/docs/rdoc?l=en;v=14.7;d=ug612.pdf (http://www.xilinx.com/cgi-bin/docs/rdoc?l=en;v=14.7;d=ug612.pdf)
-
- Trivial changes to some VHDL signals, which clearly have nothing to do with the critical paths, also can throw the result off track significantly.
* I switch the clock generation from PLL (which is power-hungry) to DCM_CLKGEN, which makes the clock jitter a mere 25ps worse. But the penalty on the resulting overall design clock speed is nearly 10 MHz?! Both clock generators are directly driving a BUFG, so I would assume clock distribution behaves the same for both of them.
Trivial changes re-seed how the placer starts. This is why you might see the failing paths change on each run. Changing from PLL to DCM does that, too, even if the source code otherwise is unchanged.
What's the right way of doing this? I have not found any good tutorials with specific advice on how to use constraints in a robust way. Thanks for any hints you might have!
Awhile ago I did an S6 design at 100 MHz which used a lot of block RAMs, and I don't remember any particular timing challenges.
So hints. First, set the period constraint to 100 MHz. Setting it higher doesn't buy you anything. Setting it lower makes the tools try "less hard."
Look at synthesis results. Did the synthesis do what you expect or did it do something completely wacky? Coding style is important. Remember that if you don't write code in a way the synthesizer expects (basically it's a template matching process) then you could get poor results. The other day I looked at the synthesis report for something and saw that it inferred a boatload of flip-flops instead of the block RAM I expected. That was from a mistake in the code that inferred the memory.
-
You can read this if you haven't already: http://www.xilinx.com/cgi-bin/docs/rdoc?l=en;v=14.7;d=ug612.pdf (http://www.xilinx.com/cgi-bin/docs/rdoc?l=en;v=14.7;d=ug612.pdf)
I won't claim to have read the Timing Closure Guide, but I have read parts of it...
I didn't really find recommendations for the situation I am in. It all seems clear enough when you have "hard and fast" constraints you have to meet, e.g. relative to external signals. But in my case, the only hard constraints are that various signals need to reach the end of their longish signal paths in time before the next clock edge, and I would like to "help" the automatic optimization find the best placement and routing. Is that a good idea at all, or will it lead to even worse results when some other, unrelated placements change later? (Similar to what I found with the placement constraints?) How to go about it?
@xlnx, you may be right there -- trying to eke out the last few MHz at the edge of the possible performance may simply be an ill-defined problem. I see that shortening the signal paths by introducing intermediate registers and pipelining could provide a more robust soution than trying to nudge the optimization to the very best solution. My problem is that I can't easily do that, for two reasons:
- The design is a microprocessor softcore which is meant to run as fast as possible from on-chip BRAM. But it is also meant to run cycle-correct from external RAM/ROM (much more slowly), and switch between the two modes seamlessly. Extra pipelining in the processor seems incompatible with those requirements.
- Much simpler reason: The actual softcore is third-party code. I still hope I don't have to dig too deep, and can simply leave it as it is...
I will sleep on it and think some more about ways to improve the design while keeping the processor cycle-correct. (But have done that for a while without much success.) In the meantime, if anybody has hints on nudging the optimizer in the right direction to achieve more stable results, I'd much appreciate those!
-
Awhile ago I did an S6 design at 100 MHz which used a lot of block RAMs, and I don't remember any particular timing challenges.
No fair! ;)
Did you use them all combined into one large memory? I realize that I may be abusing the block RAM somewhat. It will certainly perform much better if it is used in localized functional modules: just few blocks of RAM interacting with logic placed nearby; and then some other blocks elsewhere interact with other logic which, again, can be placed near the RAM; etc. I assume that's what the chip designers had in mind. Long routes running half the length of the chip, from RAM at the outer edge to the CPU at the center, are my bottleneck.
So hints. First, set the period constraint to 100 MHz. Setting it higher doesn't buy you anything. Setting it lower makes the tools try "less hard."
Check. As mentioned above, I think I understand the "hard and fast" constraints: State them, but don't exaggerate.
Look at synthesis results. Did the synthesis do what you expect or did it do something completely wacky? Coding style is important. Remember that if you don't write code in a way the synthesizer expects (basically it's a template matching process) then you could get poor results. The other day I looked at the synthesis report for something and saw that it inferred a boatload of flip-flops instead of the block RAM I expected. That was from a mistake in the code that inferred the memory.
Thank you, that is a good point. I have not questioned the synthesis much, but focused on the place & route results mostly. I guess I will have to look into that processor core and understand it a bit less superficially, since that's where a lot of the action is. It is largely a largish state machine, and I have found that th synthesizer is quite clever when encoding these, but who knows...
-
Awhile ago I did an S6 design at 100 MHz which used a lot of block RAMs, and I don't remember any particular timing challenges.
No fair! ;)
Did you use them all combined into one large memory?
In a manner of speaking. The FPGA connects to eight ADC channels, each digitizing an output of an image sensor. The memory is a dual port, with a write side and a read side. (Not a "true dual port" where you can write to and read from each side independently.)
The write side looks like eight independent memories, as the output of each ADC has samples written to its associated buffer simultaneously. So say the line has 8192 pixels (and each pixel is 16 bits). Each buffer is 1024 pixels deep, so there is a ten-bit write address counter.
The read side looks like one big 8192-deep buffer, so it has a 13-bit address counter. The idea is that it de-interleaves the data "in place," so downstream functions just get lines with pixels already in order.
Actually there were two such buffers, accessed in a ping-pong fashion (writing to one while reading out of the other). This extends the size of the address counters by one bit, with the MSb being the X/Y select.
I ended up replicating the address counters for the write side, to limit the loading and the routing.
All of the memories were inferred, so no instantiation of elements from the Xilinx library.
Look at synthesis results. Did the synthesis do what you expect or did it do something completely wacky? Coding style is important. Remember that if you don't write code in a way the synthesizer expects (basically it's a template matching process) then you could get poor results. The other day I looked at the synthesis report for something and saw that it inferred a boatload of flip-flops instead of the block RAM I expected. That was from a mistake in the code that inferred the memory.
Thank you, that is a good point. I have not questioned the synthesis much, but focused on the place & route results mostly. I guess I will have to look into that processor core and understand it a bit less superficially, since that's where a lot of the action is. It is largely a largish state machine, and I have found that th synthesizer is quite clever when encoding these, but who knows...
It might just be that the processor core has long logic paths that can't be optimized. Of course if you wrote the core, or you can modify it, you can look for places where you can pipeline, but that might change how the core functions.
-
Block ram without at least 2 pipeline registers is always going to be slow, and it doesn't get much better with new FPGA families. The block rams are far away and routing delay dominates. Can you post a screenshot of the failing path highlighted in planahead (the physical device view)?
Vivado optimizes better than ISE, but if you are on spartan 6 there isn't much you can do. I never had trouble getting >200MHz on the slowest speed grade spartan 6, but everything I did was easy to pipeline.
-
Block ram without at least 2 pipeline registers is always going to be slow, and it doesn't get much better with new FPGA families. The block rams are far away and routing delay dominates.
Yes, that's exactly the bottleneck I face: The CPU softcore sits more or less in the center of the chip. (It needs a fraction of the logic cells only.) Since I need to use all 32 RAM blocks, that incudes the blocks at the far ends of the chip, and the route delays to those dominate the timing.
Vivado optimizes better than ISE, but if you are on spartan 6 there isn't much you can do.
Hence my hope that I could give ISE some "direction" via constraints, to get it to optimize somewhat better. But apparently there's not much one can do along those lines. Maybe I'll find some places for mild pipelining in the design...
It's annoying that Xilinx decided not to support Spartan-6 in Vivado. But I guess it makes sense from their perspective: They don't want Spartan-6 to be used for new designs and only keep selling it for use in existing devices. Then ISE is adequate, since all people are supposed to do with it is sustain their old designs.
-
I suggest (from the synthesis results) you take an extra look at how the BRAMs are organized. If they are organized as deep and narrow (down to 1 bit wide each), then much less multiplexing of the datalines is required, which could hugely impact the combinatorial delays.
Even if you infer all your BRAMs, take a look at how CoreGen would build your 64kB BRAM, and try to do similarly when you infer it through your own code.
-
1. Have you tried placing it into larger device just to see if maybe it's a routing congestion issue?
2. Do you have access to the source code for your CPU core? If so, adding one more pipeline stage is usually fairly trivial. Just watch out for feedback paths (signals which cross stage boundaries "the wrong way around").
3. Have you considering migrating to Spartan-7? They are significantly faster than S6, they are cheaper, and you can use Xilinx's Microblaze core for free if that's what you want.
-
I suggest (from the synthesis results) you take an extra look at how the BRAMs are organized. If they are organized as deep and narrow (down to 1 bit wide each), then much less multiplexing of the datalines is required, which could hugely impact the combinatorial delays.
I had that thought too, and tried explicitly declaring a set of eight 64k*1bit RAMs, and generating them via the IP Wizard. That did indeed win me 5 MHz initially, but again, the benefit has not been reproducible. Comes and goes with every new synthesis and place&route run...
-
1. Have you tried placing it into larger device just to see if maybe it's a routing congestion issue?
I have indeed tried to go from the LX9 to the LX16, and it brings a slight advantage. Not due to routing congestion with the LX9, it seems, but because the LX9 chip has a longish, rectangular aspect ratio, while the LX16 is more or less a square. That's how the "physical" view of the tools shows it, at least, and it seems to be real: In the LX16, the BRAM blocks at the far edges of the chip don't need to be used, and hence the maximum routes are shorter.
2. Do you have access to the source code for your CPU core? If so, adding one more pipeline stage is usually fairly trivial. Just watch out for feedback paths (signals which cross stage boundaries "the wrong way around").
Yes, it's open source. I had hoped to avoid digging into it, and more importantly, want to keep it cycle-correct when operating on the slow external bus. In principle that should be possible -- whenever an external bus access is needed, there's enough time to let the CPU catch up and empty its pipeline. Sounds a bit intricate and error-prone (on my end ;-) though...
3. Have you considering migrating to Spartan-7? They are significantly faster than S6, they are cheaper, and you can use Xilinx's Microblaze core for free if that's what you want.
Duh... no, I had not considered that yet. And I had not even realized that the S7 is lower cost than the S6; had always assumed it was the other way round by a fair margin. Better chip performance, modern toolchain, lower cost -- seems worth a look indeed, thank you!
-
Perhaps it is worth adding a few placement constraints to the design.
If only I knew where... ::)
As mentioned in the original post, I did try placement constraints. Having looked at the timing report and the critical signal paths, I had noted that some of the CPU softcore's components had been placed pretty far away from the center, adding to the routing delays. So I restricted them to a pblock in a central area.
This helped somewhat, but only initially. With some minor changes in other corners of the design, the allowable clock rate deteriorated -- and I found that removing the placement constraint made things better again, presumably because the placer found better uses for the central area. I have yet to find a placement constraint that is so strategic in nature that it will lead to robust speed improvements.
Maybe the "solution" is to keep this kind of fiddling to the very end, when nothing is expected to change anymore in the design?
-
Duh... no, I had not considered that yet. And I had not even realized that the S7 is lower cost than the S6; had always assumed it was the other way round by a fair margin. Better chip performance, modern toolchain, lower cost -- seems worth a look indeed, thank you!
It's kind of a wash at the very low end (LX4 is cheaper than S7-6, but the latter have more resources), but everything else is in favor of S7 - S7-15 is about the same price as S6-9 (the price can go either way depending on the package, speed grade, but the former has 15k "gates" vs 9k in the latter), and as you go up the ladder the difference only grows.
Give it a try as it doesn't cost you anything (except for some of your time) - import your code into Vivado and run P&R just to see what happens. One thing you've got to keep in mind though is Vivado works differently than ISE in a sense that it's timing constraints-driven - meaning it will do it's best to close the timing, but not an iota more. The reason for that I suspect is that Vivado supports gigantic devices with insane amount of resources, so there is an exponential explosion of possible configurations and consequently it will take forever to try them all, so it settles on "good enough" solution that meets all constraints. This is why proper constraints are even more important for Vivado. But the downside of that approach is that the only way to see how far can you push your design is to progressively tighten timing constraints and seeing if P&R succeeds or not. But even then there are different strategies you can try by forcing P&R to try harder to find a solution. As you get close to the limit, you will see the P&R taking longer and longer as it needs to work harder to close timing.
-
I don't think routing delays prevent you from running at 100 MHz. 100 MHz means 10ns between clock edges. This is plenty for routing.
Look at the critical path and see what's in it. Do you really see only wires, or are there any LUTs on the way? How many layers of LUTs do you have?
BRAM blocks have optional output registers. Are they used in your design? If you use these, depending on your speed grade, it may save you several ns. Of course this will introduce one clock delay, so you may need to change the design, and few MHz may not be worth it.
-
I don't think routing delays prevent you from running at 100 MHz. 100 MHz means 10ns between clock edges. This is plenty for routing.
Look at the critical path and see what's in it. Do you really see only wires, or are there any LUTs on the way? How many layers of LUTs do you have?
Believe me, I have looked long and hard... Typically 60..70% of the critical path duration is due to routing delays.
BRAM blocks have optional output registers. Are they used in your design? If you use these, depending on your speed grade, it may save you several ns. Of course this will introduce one clock delay, so you may need to change the design, and few MHz may not be worth it.
Right, that's going back to the question of cycle-correct simulation of the processor -- a 65x02, by the way. I don't use the output registers at the moment, since the additional delay cycle would significantly change the processor's behavior. It typically wants those data right away, e.g. to calculate the address for the very next bus cycle.
-
Believe me, I have looked long and hard... Typically 60..70% of the critical path duration is due to routing delays.
That's not how I would look at it. When you add LUTs, you automatically add routing to and from LUTs. Routing delays are roughly determined by the number of switch boxes on the way. Any intermediary LUT will add at least two switching boxes to the route which might be a bigger delay than the LUT itself. If you have any combinatorial logic on the way, the solution is to minimize the number of combinatorial layers, which will also remove the associated routing.
For example, you want to drive from Los Angeles to San Francisco. If you decide to visit Las Vegas on the way, your trip will get longer and you will have to drive more miles. If you want to drive less, the solution is not to find better shorter roads, but to drop Las Vegas entirely.
Although you have the unregistered BRAM (3.5 ns delay for -1 speed grade), you still have plenty of time left for routing.
-
That's not how I would look at it. When you add LUTs, you automatically add routing to and from LUTs. Routing delays are roughly determined by the number of switch boxes on the way. Any intermediary LUT will add at least two switching boxes to the route which might be a bigger delay than the LUT itself. If you have any combinatorial logic on the way, the solution is to minimize the number of combinatorial layers, which will also remove the associated routing.
For example, you want to drive from Los Angeles to San Francisco. If you decide to visit Las Vegas on the way, your trip will get longer and you will have to drive more miles. If you want to drive less, the solution is not to find better shorter roads, but to drop Las Vegas entirely.
Although you have the unregistered BRAM (3.5 ns delay for -1 speed grade), you still have plenty of time left for routing.
Well, but having the LUTs (i.e. the logic for a working CPU) is the whole point of the design, isn't it? I can't really do away with them; just moving unmodified data from BRAM to BRAM wouldn't quite cut it. ;)
I can't remove the LUTs. I could, in principle, introduce more pipelining stages, but am reluctant to do so since I don't want to mess with the original 6502 timing. So reducing the routing delays by optimal positioning of the RAMs and logic seems like an important piece of the optimization; and I had hoped that I could nudge the Xilinx tools a bit to make them find the optimum more reproducibly.
Anyway, I think I will simply settle for slightly less for now, and move on to more rewarding tasks. 96 MHz is also a nice clock rate since it aligns well with the 48 MHz USB clock... 8)
-
Well, but having the LUTs (i.e. the logic for a working CPU) is the whole point of the design, isn't it? I can't really do away with them; just moving unmodified data from BRAM to BRAM wouldn't quite cut it. ;)
Sure, but you can do a better design which does the same with less combinatorial layers. Or you can move some of the combinatorial logic from the critical path to other stages of your pipeline. There are many things you can do. Of course, if you do not want to make any changes then nothing will change ;)
You can improve timing by moving things around and doing manual placing and routing, but this is very tedious job. It would never occur to me to do this to gain 4 MHz. Get a faster speed grade, or move to S-7.
-
You can improve timing by moving things around and doing manual placing and routing, but this is very tedious job. It would never occur to me to do this to gain 4 MHz. Get a faster speed grade, or move to S-7.
Sure, the difference between 96 to 100 MHz would just be cosmetic in nature -- as I said in the OP, 100 is a three-digit number. :)
I am already using the -3 speed grade, so the XC6SLX25 (for slightly better BRAM placement) or Spartan-7 are the possible hardware upgrades. Later, maybe; for the moment I'll just swallow my pride and live with the two digits...
-
Sure, the difference between 96 to 100 MHz would just be cosmetic in nature -- as I said in the OP, 100 is a three-digit number. :)
I am already using the -3 speed grade, so the XC6SLX25 (for slightly better BRAM placement) or Spartan-7 are the possible hardware upgrades. Later, maybe; for the moment I'll just swallow my pride and live with the two digits...
In practical terms, if it closes at 96 MHz, it'll work at 100 MHz too. Just cheat it by saying your input clock is 24 MHz, but in reality feed it 25 MHz, or something like that. Chances that 4% speed increase kills it are slim to none.
-
I don't think routing delays prevent you from running at 100 MHz. 100 MHz means 10ns between clock edges. This is plenty for routing.
Uh yeah. I have a design on a Spartan 6 that is not even that complex (but it's a CPU). I get a 110MHz max frequency, and among the longest paths, routing delay alone takes up to 7ns. My design uses 23% of the FPGA resources, so it's far from congested. I use almost all block RAM available though.
-
Agreed. I've run into timing issues at 125MHz on a Zynq 7014S with 30% utilisation. Lots of AXI buses though.
-
Duh... no, I had not considered that yet. And I had not even realized that the S7 is lower cost than the S6; had always assumed it was the other way round by a fair margin. Better chip performance, modern toolchain, lower cost -- seems worth a look indeed, thank you!
It's kind of a wash at the very low end (LX4 is cheaper than S7-6, but the latter have more resources), but everything else is in favor of S7 - S7-15 is about the same price as S6-9 (the price can go either way depending on the package, speed grade, but the former has 15k "gates" vs 9k in the latter), and as you go up the ladder the difference only grows.
I'm super interested to hear that, I'd been sticking with S6 and ISE in a VM just down to cost.
Current pricing on the S6 - 16k is <$6USD Qty 1.
https://lcsc.com/product-detail/CPLD-FPGA_XILINX_XC6SLX16-2FTG256C_XILINX-XC6SLX16-2FTG256C_C39313.html
Where can we get this sort of pricing on S7?
-
I'm super interested to hear that, I'd been sticking with S6 and ISE in a VM just down to cost.
Current pricing on the S6 - 16k is <$6USD Qty 1.
https://lcsc.com/product-detail/CPLD-FPGA_XILINX_XC6SLX16-2FTG256C_XILINX-XC6SLX16-2FTG256C_C39313.html
Where can we get this sort of pricing on S7?
I was comparing prices at DK and Mouser. You might want to talk to Xilinx sales to see if you can get a better deal (I'm sure you will, the only question is just how good will it be). Though smaller S7 are only available in packages with 100 user IO pins (FTGB196 with 1 mm pitch, CSGA225 with 0.8 mm pitch, and CPGA196 with 0.5 mm pitch), if you want more than that in 1 mm pitch package, you will have to step up to S50 and FGGA484 package (it has 250 IO pins).
-
There are many things you can do. Of course, if you do not want to make any changes then nothing will change ;)
I think this is where the discussion fails. There are many possible ways to improve timing on a design, but they usually involve significant architectural changes to the HDL and not just a sprinkling of magic pixie dust in the constraints/build parameters.
1) learn what the critical path is (if it radically changes from build to build thats a hint you're resource constrained, see 3b)
2) redesign the structure or clocking of the critical path to improve timing along it
2b) add constraints to make the tools find a quick solution on the critical path (based on some obvious mistake that is visible in routing at step 1)
3) return to 1)
3b) diversify the build flow (ISE smart explorer) and hope for a lucky solution among many many attempts
Most of this is best left for once the design is "finished" and ready to package up as a release, trying to chase the last 10% of timing in the middle of the design is usually a waste of time.
-
Uh yeah. I have a design on a Spartan 6 that is not even that complex (but it's a CPU). I get a 110MHz max frequency, and among the longest paths, routing delay alone takes up to 7ns. My design uses 23% of the FPGA resources, so it's far from congested. I use almost all block RAM available though.
I have explained this few posts back (even with geographical examples :) ). You have many logic layers. The logic layers are connected through switch boxes. These switch boxes produce the routing delay you see. If you divide the routing delay between logic layers, and you'll see that a single logic layer doesn't need big routing delay. But when you have many layers the routing delay grows. This has nothing to do with covering distances within FPGA, or congestion, or anything of that sort. Decrease the number of logic layers, and the routing between them will go away. That's how you make the design work faster.
-
All -- I just wanted to say thank you again for many helpful hints in this thread! I have understood that trying to optimize for the last few percent of performance at the very edge of the design is not a good idea, and is not what the timing and placement constraints are meant to achieve.
Since I am so close to my symbolic 100 MHz target, I will go for the pragmatic "solution" suggested by NorthGuy and tggzz, and just overclock by a few percent. This is a hobby project anyway, nothing mission-critical.
As a next step, I will look into Spartan-7 FPGAs. (Although in this particular application they do end up being 50% more expensive than the Spartan-6, even at Mouer prices for both, since my BRAM needs require a larger S7 chip.) But I am certainly tempted to switch to Vivado, and see how the S6 and S7 compare performance-wise.
-
I really like the Zynq. Once you get used to some of its idiosyncrasies, it is a *really* nice platform. You can easily move GB/s of data around. The platform is pretty easy to understand. The A9 is a fast ARM, even though the clock rate is not particularly high.
Vivado is f**king terrible though.
-
Uh yeah. I have a design on a Spartan 6 that is not even that complex (but it's a CPU). I get a 110MHz max frequency, and among the longest paths, routing delay alone takes up to 7ns. My design uses 23% of the FPGA resources, so it's far from congested. I use almost all block RAM available though.
I have explained this few posts back (even with geographical examples :) ). You have many logic layers. The logic layers are connected through switch boxes. These switch boxes produce the routing delay you see. If you divide the routing delay between logic layers, and you'll see that a single logic layer doesn't need big routing delay. But when you have many layers the routing delay grows. This has nothing to do with covering distances within FPGA, or congestion, or anything of that sort. Decrease the number of logic layers, and the routing between them will go away. That's how you make the design work faster.
Well, true in general. (But I still consider that route delay. Xilinx tools do as well.)
In one of my examples, I get : Data Path Delay: 8.611ns (Levels of Logic = 5)
(2.031ns logic, 6.580ns route)
5 levels is significant, but not that large. But on this path, there is some Block RAM.
There are actually paths with a slightly larger number of logic levels and shorter route delay.
That may be "fixable" by helping placement I guess, but I didn't bother as 100 MHz was my goal here.
-
5 levels is significant, but not that large. But on this path, there is some Block RAM.
There are actually paths with a slightly larger number of logic levels and shorter route delay.
When you start from the BRAM without output register, there will be big initial delay within BRAM, which Xilinx consider a routing delay too.
Regardless, if you have a chain of elements to connect there are delays associated with the elements - internal delays within the element (BEL delays as they call it), launching delays (to reach the switch box), pin delays (to get the signal from switch box to the element). Plus every element will connect to two fixed switch boxes - one for the input and one on the output. The switches within boxes will produce further delays. Whether you consider these to be routing delays (as Xilinx does) or not, you must endure all of these delays as soon as you add the element to your chain. These cannot be eliminated by better routing (althogh you can improve some of these, for example by selecting faster LUT pins). If you've got 10 ns of these, you cannot run at 100 MHz, no matter how good your routing is. It is therefore outright counterproductive to try to solve the problem by improving routing.
However, you can get in a situation where you know the solution exists, but the tools cannot find it. This can be caused by either congestion, or by high clock speed. You can try to fix the congestion problem by using better floorplanning, and you may succeed. Similarly, if you have high speed project, you can manually place your elements, sometimes duplicate some logic, sometimes route manually. These are routing problems - you have acceptable design, but it will only work if the routing improves.
And, of course, you need to be able to distingush one from the other.
-
I have a nearly completed VHDL design for the Spartan 6, implemented via ISE 14.7. The automated synthesis and place/route process reliably produces a design which runs at above 90 MHz, but then things get erratic. (Yep, 90 MHz are not a whole lot -- the major bottleneck is that I am using all block RAM as a large 64 kByte RAM, so there are long routing delays to and from the outer blocks.)
I would like to push this to 100 MHz, mainly because that's a three-digit number. ;)
I can come close, and did get timing closure at 100 MHz once, but the optimization results have been frustrating. It seems pretty clear that the place & route does not always find the optimal result. E.g.:
- I set the internal clock to 94 MHz, and the timing results tell me that my design can run at up to 98.x MHz. So I set the clock to 98 MHz -- and get a design that only runs at up to 92 MHz. I did check the clock jitter and it did not increase with the adjustment to 98 MHz, so that's not the cause.
- Trivial changes to some VHDL signals, which clearly have nothing to do with the critical paths, also can throw the result off track significantly.
- I switch the clock generation from PLL (which is power-hungry) to DCM_CLKGEN, which makes the clock jitter a mere 25ps worse. But the penalty on the resulting overall design clock speed is nearly 10 MHz?! Both clock generators are directly driving a BUFG, so I would assume clock distribution behaves the same for both of them.
Timing and placement constraints are probably the answer to my frustration, but I struggle with them. Based on earlier advice here, I have declared some false paths for signals derived from stationary inputs, and at least that seems to do no harm --- although I am not sure about reliable improvements. I have also tried placement constraints to put some critical functionality fed by the BRAM close to the chip's center, to ensure it's equidistant to the outermost BRAMs on both ends of the chip. This helps in one particular place/route run, but then when I change some other trivial detail, the placement restriction actually makes the results worse compared to a run where I remove it.
What's the right way of doing this? I have not found any good tutorials with specific advice on how to use constraints in a robust way. Thanks for any hints you might have!
It is a bit of a black art to get a design routed and close timing. First of all I'd start by setting the clock constraint to 100Mhz and don't mess with lower frequencies. Make sure to have timing constraints from the inputs to the flipflops (PADS(*) to all_ffs) and from the flipflops to the pads (all_ffs to PADS(*)).
Then route your design. If it fails use the timing analyser to see which paths are the slowest ones. It may help to change the setting to have more paths in the timing analyser log so more failing paths can be checked. If you only need a little bit of extra speed you can try to place an element manually and/or play a bit with the 'starting placer cost table'. Something else to look at is how full your FPGA is. If it isn't very full then disable the optimisation. This sounds crazy but I have a reasonably large project which uses an oversized FPGA. If I enable optimisation all the logic gets packed together and the timing towards the IO pins is impossible to achieve.
-
3. Have you considering migrating to Spartan-7? They are significantly faster than S6, they are cheaper, and you can use Xilinx's Microblaze core for free if that's what you want.
Duh... no, I had not considered that yet. And I had not even realized that the S7 is lower cost than the S6; had always assumed it was the other way round by a fair margin. Better chip performance, modern toolchain, lower cost -- seems worth a look indeed, thank you!
This is not always true, and newer is not always better. I have a project using the S6 because it is just as cheap, or cheaper than the equivalent S7, but also mostly because the 7-series is not hobbyist friendly. The newer chips are pushing the packages smaller and smaller, and a 0.5mm ball-pitch CSP (or smaller) is just not doable for hobbyists yet. The PCB features needed to route and escape the chip-scale stuff are 6-layers minimum and 0.0762mm (3mil) traces. The "standard" affordable boards are 4-layer and 0.127mm (5mil) at best, with 0.25mm (10mil) drills for vias.
edit: I stand corrected, the S7 has just about the same size and ball-pitch packages as the S6 in the low-end of the chip range.
For me, the S7 did not offer the right size package + features + price that was right, but I could get the S6 in exactly what I needed (15mm @ 0.8mm ball pitch), it was available at DigiKey and Mouser, it was cheap ($20 per chip), and I can manage to assemble the prototypes in my reflow oven. I could not go to a larger package because I'm space constrained by a PCB size requirement.
edit: The S7 has the same package, but only half the available physical I/O in that package and price.
Also, IMO, Vivado is awful compared to ISE. I have a few 7-series devboards, and the first time I went to use Vivado, I thought I had done something wrong because my "blink an LED" design took, literally, 10-minutes to generate a bit-stream. I could not believe it. And Vivado is a huge bloated mess. At lease ISE is smaller-ish, and works pretty well.
edit: I have been "informed" I am doing something wrong and I should get on the bandwagon.
Yes, it is unfortunate and really crappy the way Xilinx handled the transition from ISE to Vivado, right on the cusp of Win7 to Win10, and dropping support for S6 in Vivado. They lost some of my respect for sure, but clearly they do not care about that. Also, I managed to find (somewhere) on the Xilinx website that the S6 line is going to be in production through 2024 or 2026 (can't remember exactly), and that was enough for my project.
Sorry for the rant, it just gets old when people always push the latest thing, for the sake of it being the latest. But sometimes you cannot always jump on the latest and greatest, and sometimes the old and reliable is good enough (or better).
As for the timing problem with BRAM, it seems you have gotten a lot of help and I hope something worked for you. I have used a S6 with a design at 100MHz that used all the BRAM, and did not have too much trouble. The 100MHz design even came from a Spartan-3E, but the largest single block was only 16K in that case. All I can suggest is to register your inputs and outputs, and add a pipe-line (register) stage to help ease the routing constraints.
-
I'm seriously impressed by the amount of crap you managed to stuff in a single post.
This is not always true, and newer is not always better. I have a project using the S6 because it is just as cheap, or cheaper than the equivalent S7, but also mostly because the 7-series is not hobbyist friendly. The newer chips are pushing the packages smaller and smaller, and a 0.5mm ball-pitch CSP (or smaller) is just not doable for hobbyists yet. The PCB features needed to route and escape the chip-scale stuff are 6-layers minimum and 0.0762mm (3mil) traces. The "standard" affordable boards are 4-layer and 0.127mm (5mil) at best, with 0.25mm (10mil) drills for vias.
This is pure BS. S7 has just one package with 0.5 mm pitch, all other packages are 0.8 and 1 mm pitch. All FPGAs I used so far were in 1 mm pitch packages.
So - year - newer isn't always better, but in this case it is better. MUCH better.
For me, the S7 did not offer the right size package + features + price that was right, but I could get the S6 in exactly what I needed (15mm @ 0.8mm ball pitch), it was available at DigiKey and Mouser, it was cheap ($20 per chip), and I can manage to assemble the prototypes in my reflow oven. I could not go to a larger package because I'm space constrained by a PCB size requirement.
Well - guess what - S7 comes in the very same 15 mm @ 0.8 mm pitch package (CSGA324), there is also smaller package 13 mm @ 0.8 mm pitch available (CSGA225). But my favorite one is FTBG196 - 15 mm @ 1 mm pitch with 100 user IO and fully routable on any 4 layer process. That just shows me that your didn't bother doing any research.
Also, IMO, Vivado is awful compared to ISE. I have a few 7-series devboards, and the first time I went to use Vivado, I thought I had done something wrong because my "blink an LED" design took, literally, 10-minutes to generate a bit-stream. I could not believe it. And Vivado is a huge bloated mess. At lease ISE is smaller-ish, and works pretty well.
That's one more sign that you've done something wrong. Vivado is waay ahead of ISE, features-wise ISE is stone age.
Yes, it is unfortunate and really crappy the way Xilinx handled the transition from ISE to Vivado, right on the cusp of Win7 to Win10, and dropping support for S6 in Vivado. They lost some of my respect for sure, but clearly they do not care about that. Also, I managed to find (somewhere) on the Xilinx website that the S6 line is going to be in production through 2024 or 2026 (can't remember exactly), and that was enough for my project.
Do you use VHDL for your designs? It fits the pattern.
-
I'm seriously impressed by the amount of crap you managed to stuff in a single post.
I realize that you were not talking to me, but since I am also pondering the Spartan 6 vs. 7 decision (triggered by your earlier suggestion -- thanks again for that!), your post does make me wonder. Why the aggressive tone? You seem to be "emotionally invested" in the Spartan-7 line? ::)
I agree with you that the available packages are not a reason to stay away from the S7. The fact that Xilinx has retained the 0.8mm pitch BGAs is a plus for me; those are nicely usable with the low-cost PCBs available from JLCPCB, Aisler etc.; other FPGA brands don't seem to offer that format.
But I do see other disadvantages vs. the S6 in my current project -- which has some space and power constraints, and needs 64 kByte of on-chip RAM: There seems to be no way around a third power rail for the S7 (1.0V Core, 1.8V Aux, 2.5V or 3.3V I/O). Modelled power consumption for the S7 is actually higher than in the equivalent S6 design, due to its 60mW "static" base load and the power-hungry clock managers. And the smallest S7 which meets my memory needs is the XC7S25, which is 50% more expensive than the XC6SLX9.
So it's clearly "horses for courses", and your "newer is better" statement oversimplifies things.
Do you use VHDL for your designs? It fits the pattern.
Now what's that supposed to mean? Why don't you toss in a "Linux vs. Windows" debate as well, for good measure? ???
-
I realize that you were not talking to me, but since I am also pondering the Spartan 6 vs. 7 decision (triggered by your earlier suggestion -- thanks again for that!), your post does make me wonder. Why the aggressive tone? You seem to be "emotionally invested" in the Spartan-7 line? ::)
I'm emotionally invested in preventing people from spreading BS. We have too much of it already.
I agree with you that the available packages are not a reason to stay away from the S7. The fact that Xilinx has retained the 0.8mm pitch BGAs is a plus for me; those are nicely usable with the low-cost PCBs available from JLCPCB, Aisler etc.; other FPGA brands don't seem to offer that format.
Yea, that's the part I really like. Even super-high pin count packages retain 1 mm pitch, which is very rare nowadays.
But I do see other disadvantages vs. the S6 in my current project -- which has some space and power constraints, and needs 64 kByte of on-chip RAM: There seems to be no way around a third power rail for the S7 (1.0V Core, 1.8V Aux, 2.5V or 3.3V I/O). Modelled power consumption for the S7 is actually higher than in the equivalent S6 design, due to its 60mW "static" base load and the power-hungry clock managers. And the smallest S7 which meets my memory needs is the XC7S25, which is 50% more expensive than the XC6SLX9.
Take a look at Artix-15. It's got more memory, though it's probably still more expensive than XC6SLX9. I personally don't care for such low end (I use parts mostly in 35-100K cells range), so not sure.
As for power rails - S7 does not require 3.3 V rail, so you can get away with just two - 1.0 V for Vccint and 1.8 V for PLL and IO. 1.8 V rail can be reused for DDR2 memory (which is what the project in my signature does), so it's hardly a waste. And the only disadvantage of DDR2 over DDR3 is this case is capacity (because S7 can only go up to 400 MHz, which is achievable for both DDR2 and DDR3). Though for low power you might want to choose DDR3L. I'm not sure about that because none of my designs are battery-powered, so my only concern as far as power goes is thermal - basically making sure nothing gets too hot.
Also you seem to ignore the fact that 7 series fabric and hard IP blocks are two times faster than in S6. That's got to count for something I think.
Now what's that supposed to mean? Why don't you toss in a "Linux vs. Windows" debate as well, for good measure? ???
No, I explained it above - it fits the pattern.
-
Take a look at Artix-15. It's got more memory, though it's probably still more expensive than XC6SLX9.
Similar price as the Spartan-25. And unfortunately the Artix series is not available in the 225-ball, 13*13 mm²BGA.
As for power rails - S7 does not require 3.3 V rail, so you can get away with just two - 1.0 V for Vccint and 1.8 V for PLL and IO.
Right -- I neglected to mention that I need at least 2.5V for the I/Os. For the S6, that means 1.2V core and 2.5V or 3.3V combined Aux/IO voltage; for the S7 I need three supplies. Again, the choice of the "right" FPGA depends on the task at hand.
Also you seem to ignore the fact that 7 series fabric and hard IP blocks are two times faster than in S6. That's got to count for something I think.
That might count for something. I need to come up to speed with Vivado and actually try to put my design on the Spartan-7. Looking at the S6 and S7 datasheets, I surprisingly couldn't find the claimed higher speed of the S7. The stated BRAM and CLB access times were actually slower in the S7. I might have overlooked other aspects there, of course; the proof will be in generating the design.
No, I explained it above - [using VHDL] fits the pattern.
I still am not sure I get your drift. If you meant to imply that using VHDL fits a pattern of producing BS, I hereby feel officially insulted. :-\
(I, for one, use VHDL almost exclusively. I get annoyed with its wordiness at times, but it has easily made up for that with all the errors of mine which it has found already during synthesis. And I am not aware of any difference in the synthesis results. So again, what's your point?)
-
The stated BRAM and CLB access times were actually slower in the S7.
This is not accurate. Looking at the docs (DS162 vs DS189).
Max BRAM frequency is 280 MHz for S6-2 and 388 MHz for S7-2.
The combinatory LUT delay (for the fastest input pin) is 0.26 ns for S6-2 vs 0.11 ns for S7-2.
S7 is exactly the same as A7, except there's no -3 grade.
-
Similar price as the Spartan-25. And unfortunately the Artix series is not available in the 225-ball, 13*13 mm²BGA.
If you can make do with a bit over 100 user IO, you might want to take a look at CPG236 or CPG238 packages. They are 0.5 mm pitch, but if you will look at the pinout diagram (see attachment, it shows power pins), you will see that you can fully route it out on just 2 layers (provided you can fit a trace between 0.5 mm pads).
Right -- I neglected to mention that I need at least 2.5V for the I/Os. For the S6, that means 1.2V core and 2.5V or 3.3V combined Aux/IO voltage; for the S7 I need three supplies. Again, the choice of the "right" FPGA depends on the task at hand.
That is not FPGA requirement, If my design requires 5 different power rails, I can't hold it against FPGA. But again - the current trend is towards lowering interface voltage, so I appreciate that they eliminated 2.5 V rail which is very rarely used (except for LVDS) - typically it's either 3.3 or 1.8 V. So I will choose 1.8 V over 2.5 V any time, because it's actually useful (just about all my boards have some sort of 1.8 V peripherals, so that rail would be required anyway).
That might count for something. I need to come up to speed with Vivado and actually try to put my design on the Spartan-7. Looking at the S6 and S7 datasheets, I surprisingly couldn't find the claimed higher speed of the S7. The stated BRAM and CLB access times were actually slower in the S7. I might have overlooked other aspects there, of course; the proof will be in generating the design.
It only takes a single look at datasheets to see the difference. S6 BRAM Fmax is 320 MHz for the fastest speed grade, but in S7 Fmax for BRAM is 388 MHz for the *slowest* speed grade. The difference is even more impressive for DSP tiles, despite the fact that S7 DSP is more advanced than in S6. BRAM's clock-to-out time is about two times better if you use output register (which you should if you care about performance).
I still am not sure I get your drift. If you meant to imply that using VHDL fits a pattern of producing BS, I hereby feel officially insulted. :-\
(I, for one, use VHDL almost exclusively. I get annoyed with its wordiness at times, but it has easily made up for that with all the errors of mine which it has found already during synthesis. And I am not aware of any difference in the synthesis results. So again, what's your point?)
There is a bit of a history on this forum, when few avid VHDL propagandists are content to spread all kind of BS without bothering to do any research first to see if their arguments actually have any merit. It doesn't have anything to do with you, unless you are one of them.
-
The stated BRAM and CLB access times were actually slower in the S7.
This is not accurate. Looking at the docs (DS162 vs DS189).
Max BRAM frequency is 280 MHz for S6-2 and 388 MHz for S7-2.
The combinatory LUT delay (for the fastest input pin) is 0.26 ns for S6-2 vs 0.11 ns for S7-2.
S7 is exactly the same as A7, except there's no -3 grade.
I was referring to BRAM clock to data out time, which is critical in my design (1.85ns for S6, 2.13ns for S7) and to CLB AX..DX to output times. I may have compared apples to oranges when looking at the CLB times; the slower BRAM access time is clearly relevant to me.
Btw, I did compare the fastest speed grades for both families, i.e. -3 for the Spartan 6, as I did when comparing prices.
-
I was referring to BRAM clock to data out time, which is critical in my design (1.85ns for S6, 2.13ns for S7) and to CLB AX..DX to output times. I may have compared apples to oranges when looking at the CLB times; the slower BRAM access time is clearly relevant to me.
You can use BRAM with output register, or without. The path from output register is made much faster in 7-series, so you can achieve much faster reading frequency.
The AX .. DX are inputs for dedicated MUXes, which are slow anyway. You probably will never see them on your critical path.
BTW, LUT delays are sort of phony anyway. The datasheet lists 0.11 ns for LUT for S7-2, but this is only for the fastest pin. For all other pins there are pin delays which dwarf the number. For example, pin delays from the S7-2 speed file are:
A1 - 0.435 ns
A2 - 0.428 ns
A3 - 0.283 ns
A4 - 0.231 ns
A5 - 0.090 ns
A6 - 0.004 ns
When this get combined with the LUT delays, you get real numbers. Say, pin A1 will have 0.54 ns delay, nothing close to 0.11 ns for pin A6. That's the same pattern for S-6. So, there's no reason to dwell on posted numbers - these are for marketing. Look at your design closure instead. I can tell you, 7-series are faster.
-
I'm seriously impressed by the amount of crap you managed to stuff in a single post.
I'm seriously impressed by the amount of rudeness you managed to stuff into a single reply.
This is pure BS. S7 has just one package with 0.5 mm pitch, all other packages are 0.8 and 1 mm pitch. All FPGAs I used so far were in 1 mm pitch packages.
So - year - newer isn't always better, but in this case it is better. MUCH better.
This is pure BS, that you need to be such a jerk about correcting someone. Not everyone is unwilling to admit they made a mistake. Sorry for being human.
It is true, and I stand corrected, the 7S15 does have the same packaging I am using. I had to go review, since I looked hard and long at the S7 when I was making a choice. After reviewing again, I see the 7S15 in the CSG225 package only has 100 I/Os over two adjacent banks. Cost is about $20, so similar to the LX9 I'm using. It is not until you get into the 7S25 that you get the extra I/O, but the 7S25 is in the $35 to $40 range.
[attach=1]
Well - guess what - S7 comes in the very same 15 mm @ 0.8 mm pitch package (CSGA324), there is also smaller package 13 mm @ 0.8 mm pitch available (CSGA225). But my favorite one is FTBG196 - 15 mm @ 1 mm pitch with 100 user IO and fully routable on any 4 layer process. That just shows me that your didn't bother doing any research.
Let me offer an alternative: I did my research over a year ago and did not remember exactly why I had to rule out the S7. I thought it was the packaging because I really wanted to use it. I apologize for making the mistake of not remembering correctly, and I certainly should have reviewed the S7 packages before posting. Your reply just shows me that you must be an intolerant person, and no one should make mistakes around you.
That's one more sign that you've done something wrong. Vivado is waay ahead of ISE, features-wise ISE is stone age.
Clearly I'm a walking mistake and better keep my mouth shut. Thanks for making that clear. Never mind trying to suggest looking at Vivado again because maybe it has gotten a lot better since I tried it last, or anything else constructive, to keep the conversation helpful.
Do you use VHDL for your designs? It fits the pattern.
Do you use Verilog or SV for your designs? It fits the pattern of your attitude.
@ebastler: I apologize for this, it is not what I intended or expected to happen. I was just trying to help, but I should have checked my facts first.
-
haha whenever I see asmi enter a thread I just know the pot's about to stir ;)
I'm just here avoiding the S6 vs S7 debate by using $6 salvaged Artix-7 XC7A100T ;)
-
Let me offer an alternative: I did my research over a year ago and did not remember exactly why I had to rule out the S7. I thought it was the packaging because I really wanted to use it. I apologize for making the mistake of not remembering correctly, and I certainly should have reviewed the S7 packages before posting. Your reply just shows me that you must be an intolerant person, and no one should make mistakes around you.
I usually do my research before I post anything. When I don't and write from memory, I always explicitly say so ("As far as I remember", "If my memory serves me", etc.). This way readers can see what is fact, and what is opinion (possibly wrong). I would have zero problems with your entire post if it would mention if it's opinion and not fact. I would still correct inaccuracies, but the tone would be totally different.
I certainly admire that you can admit a mistake (very few people ever do that, and even less so online), and I do apologize for a bit of a harsh tone - it's often hard to "feel" the person behind the text.
-
I'm just here avoiding the S6 vs S7 debate by using $6 salvaged Artix-7 XC7A100T ;)
For commercial production? :o
-
For commercial production? :o
Nah just hobby projects. The only commercial project I've been involved with so far that used an FPGA used a cheap spartan 6 (XC6SLX9).
-
Nah just hobby projects. The only commercial project I've been involved with so far that used an FPGA used a cheap spartan 6 (XC6SLX9).
Well we're discussing real commercial products here. The calculus changes quite a bit when you start shipping your stuff to the customers and you are required to provide support in case something happens in the field.
For commercial prototypes the price of chips is not important as it's still peanuts in R&D budget, and majority of it goes towards paying for engineers' time. When you have to pay 600$ or more per day to every single engineer employed in the project, even 500$+ chips are little more than a pocket change in the grand scheme of things.
-
Well we're discussing real commercial products here.
Oops, I didn't realize that.
Pure amateur here. May I stay around? ::)
-
Oops, I didn't realize that.
Pure amateur here. May I stay around? ::)
Of course you can. Just keep in mind that nobody should ever buy Xilinx chips from the likes of DK for production. If you talk to Xilinx directly, you can get much better prices even if your volume is not very high.
-
Oops, I didn't realize that.
Pure amateur here. May I stay around? ::)
Of course you can. Just keep in mind that nobody should ever buy Xilinx chips from the likes of DK for production. If you talk to Xilinx directly, you can get much better prices even if your volume is not very high.
Uhhh, we (day job) buy Xilinx chips from DigiKey for production. Why? Because we aren't buying the chips by the pallet-load. Xilinx doesn't want to deal with us.
And the price of the chips is a pittance compared to the price of the products. In fact, we "standardize" on two variants of the Artix-7 (and before that, S6) and use them even if they are very oversized for the design, simply to get the price breaks for quantity purchases. And that quantity still isn't enough for Xilinx to care.
-
Uhhh, we (day job) buy Xilinx chips from DigiKey for production. Why? Because we aren't buying the chips by the pallet-load. Xilinx doesn't want to deal with us.
And the price of the chips is a pittance compared to the price of the products. In fact, we "standardize" on two variants of the Artix-7 (and before that, S6) and use them even if they are very oversized for the design, simply to get the price breaks for quantity purchases. And that quantity still isn't enough for Xilinx to care.
Why not? Single pallet is only 90 chips. If even that is way too much, then it's not really that kind of production that I was talking about. I'd say it's closer to one-offs, and low volume always have higher overhead.
-
Uhhh, we (day job) buy Xilinx chips from DigiKey for production. Why? Because we aren't buying the chips by the pallet-load. Xilinx doesn't want to deal with us.
And the price of the chips is a pittance compared to the price of the products. In fact, we "standardize" on two variants of the Artix-7 (and before that, S6) and use them even if they are very oversized for the design, simply to get the price breaks for quantity purchases. And that quantity still isn't enough for Xilinx to care.
Why not? Single pallet is only 90 chips. If even that is way too much, then it's not really that kind of production that I was talking about. I'd say it's closer to one-offs, and low volume always have higher overhead.
Your definition of pallet differs from mine (https://en.wikipedia.org/wiki/Pallet). So when I say "pallet," I mean "Lots of stuff loaded onto a pallet so the pile is about five feet high, and the whole deal is strapped down, and a forklift is used to load the deal onto the truck."
A pallet is not 90 chips. We call that a "tray." Which we buy.
Yes, your production greatly differs from ours. And my point, which you missed, was that you used the phrase "nobody should ever buy Xilinx chips from the likes of DK for production," without specifying quantity of production. That we do small-scale production of specialist products doesn't mean we don't do production.
-
Your definition of pallet differs from mine (https://en.wikipedia.org/wiki/Pallet). So when I say "pallet," I mean "Lots of stuff loaded onto a pallet so the pile is about five feet high, and the whole deal is strapped down, and a forklift is used to load the deal onto the truck."
A pallet is not 90 chips. We call that a "tray." Which we buy.
Right. I meant to say "tray". I was told you can buy a single tray directly from Xilinx, though haven't tried it yet, but planning to do so in the near future.
Yes, your production greatly differs from ours. And my point, which you missed, was that you used the phrase "nobody should ever buy Xilinx chips from the likes of DK for production," without specifying quantity of production. That we do small-scale production of specialist products doesn't mean we don't do production.
I'm doing the same thing right now, but I don't really consider it "production" in a sense it's not "mass production". Of course it still is production in a project sense (meaning the boards actually go to customers), I should've been more precise with terminology. I'm sorry for confusion.
-
Your definition of pallet differs from mine (https://en.wikipedia.org/wiki/Pallet). So when I say "pallet," I mean "Lots of stuff loaded onto a pallet so the pile is about five feet high, and the whole deal is strapped down, and a forklift is used to load the deal onto the truck."
A pallet is not 90 chips. We call that a "tray." Which we buy.
Right. I meant to say "tray". I was told you can buy a single tray directly from Xilinx, though haven't tried it yet, but planning to do so in the near future.
Yes, your production greatly differs from ours. And my point, which you missed, was that you used the phrase "nobody should ever buy Xilinx chips from the likes of DK for production," without specifying quantity of production. That we do small-scale production of specialist products doesn't mean we don't do production.
I'm doing the same thing right now, but I don't really consider it "production" in a sense it's not "mass production". Of course it still is production in a project sense (meaning the boards actually go to customers), I should've been more precise with terminology. I'm sorry for confusion.
Apology accepted. Even from a Verilog fan.
-
Even from a Verilog fan.
You confuse me with someone else. I'm not a Verilog fan.