Author Topic: Xilinx ISE - how to improve erratic timing optimization?  (Read 19930 times)

0 Members and 1 Guest are viewing this topic.

Offline ebastlerTopic starter

  • Super Contributor
  • ***
  • Posts: 7763
  • Country: de
Xilinx ISE - how to improve erratic timing optimization?
« on: June 13, 2020, 09:04:49 am »
I have a nearly completed VHDL design for the Spartan 6, implemented via ISE 14.7. The automated synthesis and place/route process reliably produces a design which runs at above 90 MHz, but then things get erratic. (Yep, 90 MHz are not a whole lot -- the major bottleneck is that I am using all block RAM as a large 64 kByte RAM, so there are long routing delays to and from the outer blocks.)

I would like to push this to 100 MHz, mainly because that's a three-digit number. ;)
I can come close, and did get timing closure at 100 MHz once, but the optimization results have been frustrating. It seems pretty clear that the place & route does not always find the optimal result. E.g.:
  • I set the internal clock to 94 MHz, and the timing results tell me that my design can run at up to 98.x MHz. So I set the clock to 98 MHz -- and get a design that only runs at up to 92 MHz. I did check the clock jitter and it did not increase with the adjustment to 98 MHz, so that's not the cause.
  • Trivial changes to some VHDL signals, which clearly have nothing to do with the critical paths, also can throw the result off track significantly.
  • I switch the clock generation from PLL (which is power-hungry) to DCM_CLKGEN, which makes the clock jitter a mere 25ps worse. But the penalty on the resulting overall design clock speed is nearly 10 MHz?! Both clock generators are directly driving a BUFG, so I would assume clock distribution behaves the same for both of them.
Timing and placement constraints are probably the answer to my frustration, but I struggle with them. Based on earlier advice here, I have declared some false paths for signals derived from stationary inputs, and at least that seems to do no harm --- although I am not sure about reliable improvements. I have also tried placement constraints to put some critical functionality fed by the BRAM close to the chip's center, to ensure it's equidistant to the outermost BRAMs on both ends of the chip. This helps in one particular place/route run, but then when I change some other trivial detail, the placement restriction actually makes the results worse compared to a run where I remove it.

What's the right way of doing this? I have not found any good tutorials with specific advice on how to use constraints in a robust way. Thanks for any hints you might have!
 

Offline xlnx

  • Contributor
  • Posts: 20
  • Country: no
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #1 on: June 13, 2020, 09:57:55 am »
What you describe is what happens when you are on the edge of what performance you can expect from the device, given your design. When you turn up the clock to 98MHz, the tool will give up faster when it figures out it cannot be done, or it has already tried different strategies that turned out to be worse.

Your frustrations with this will not end unless you make changes in the code itself. It is often a question of low latency vs. high bandwidth - and it seems that you have already chosen low latency for fast access to your BRAM. And it will only get worse when you start filling up the device with more features.

You could a) pay more money for a higher speed grade device, or b) Tweak your design towards higher bandwidth, by adding more registers/pipelining in your critical paths. Also consider utilizing full-bandwidth elastic buffers (double-registers or SRL16's) to build tiny FIFOs which consumes very litt resources.

 

Online SiliconWizard

  • Super Contributor
  • ***
  • Posts: 17795
  • Country: fr
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #2 on: June 13, 2020, 04:32:07 pm »
You can read this if you haven't already: http://www.xilinx.com/cgi-bin/docs/rdoc?l=en;v=14.7;d=ug612.pdf
 

Offline Bassman59

  • Super Contributor
  • ***
  • Posts: 2499
  • Country: us
  • Yes, I do this for a living
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #3 on: June 13, 2020, 06:54:09 pm »
  • Trivial changes to some VHDL signals, which clearly have nothing to do with the critical paths, also can throw the result off track significantly.

    * I switch the clock generation from PLL (which is power-hungry) to DCM_CLKGEN, which makes the clock jitter a mere 25ps worse. But the penalty on the resulting overall design clock speed is nearly 10 MHz?! Both clock generators are directly driving a BUFG, so I would assume clock distribution behaves the same for both of them.
Trivial changes re-seed how the placer starts. This is why you might see the failing paths change on each run. Changing from PLL to DCM does that, too, even if the source code otherwise is unchanged.

Quote
What's the right way of doing this? I have not found any good tutorials with specific advice on how to use constraints in a robust way. Thanks for any hints you might have!

Awhile ago I did an S6 design at 100 MHz which used a lot of block RAMs, and I don't remember any particular timing challenges.

So hints. First, set the period constraint to 100 MHz. Setting it higher doesn't buy you anything. Setting it lower makes the tools try "less hard."

Look at synthesis results. Did the synthesis do what you expect or did it do something completely wacky? Coding style is important. Remember that if you don't write code in a way the synthesizer expects (basically it's a template matching process) then you could get poor results. The other day I looked at the synthesis report for something and saw that it inferred a boatload of flip-flops instead of the block RAM I expected. That was from a mistake in the code that inferred the memory.
 

Offline ebastlerTopic starter

  • Super Contributor
  • ***
  • Posts: 7763
  • Country: de
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #4 on: June 13, 2020, 07:23:05 pm »
You can read this if you haven't already: http://www.xilinx.com/cgi-bin/docs/rdoc?l=en;v=14.7;d=ug612.pdf

I won't claim to have read the Timing Closure Guide, but I have read parts of it...

I didn't really find recommendations for the situation I am in. It all seems clear enough when you have "hard and fast" constraints you have to meet, e.g. relative to external signals. But in my case, the only hard constraints are that various signals need to reach the end of their longish signal paths in time before the next clock edge, and I would like to "help" the automatic optimization find the best placement and routing. Is that a good idea at all, or will it lead to even worse results when some other, unrelated placements change later? (Similar to what I found with the placement constraints?) How to go about it?

@xlnx, you may be right there -- trying to eke out the last few MHz at the edge of the possible performance may simply be an ill-defined problem. I see that shortening the signal paths by introducing intermediate registers and pipelining could provide a more robust soution than trying to nudge the optimization to the very best solution. My problem is that I can't easily do that, for two reasons:
  • The design is a microprocessor softcore which is meant to run as fast as possible from on-chip BRAM. But it is also meant to run cycle-correct from external RAM/ROM (much more slowly), and switch between the two modes seamlessly. Extra pipelining in the processor seems incompatible with those requirements.
  • Much simpler reason: The actual softcore is third-party code. I still hope I don't have to dig too deep, and can simply leave it as it is...
I will sleep on it and think some more about ways to improve the design while keeping the processor cycle-correct. (But have done that for a while without much success.) In the meantime, if anybody has hints on nudging the optimizer in the right direction to achieve more stable results, I'd much appreciate those!
 

Offline ebastlerTopic starter

  • Super Contributor
  • ***
  • Posts: 7763
  • Country: de
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #5 on: June 13, 2020, 07:36:18 pm »
Awhile ago I did an S6 design at 100 MHz which used a lot of block RAMs, and I don't remember any particular timing challenges.
No fair!  ;)

Did you use them all combined into one large memory? I realize that I may be abusing the block RAM somewhat. It will certainly perform much better if it is used in localized functional modules: just few blocks of RAM interacting with logic placed nearby; and then some other blocks elsewhere interact with other logic which, again, can be placed near the RAM; etc. I assume that's what the chip designers had in mind. Long routes running half the length of the chip, from RAM at the outer edge to the CPU at the center, are my bottleneck.

Quote
So hints. First, set the period constraint to 100 MHz. Setting it higher doesn't buy you anything. Setting it lower makes the tools try "less hard."
Check. As mentioned above, I think I understand the "hard and fast" constraints: State them, but don't exaggerate.

Quote
Look at synthesis results. Did the synthesis do what you expect or did it do something completely wacky? Coding style is important. Remember that if you don't write code in a way the synthesizer expects (basically it's a template matching process) then you could get poor results. The other day I looked at the synthesis report for something and saw that it inferred a boatload of flip-flops instead of the block RAM I expected. That was from a mistake in the code that inferred the memory.
Thank you, that is a good point. I have not questioned the synthesis much, but focused on the place & route results mostly. I guess I will have to look into that processor core and understand it a bit less superficially, since that's where a lot of the action is. It is largely a largish state machine, and I have found that th synthesizer is quite clever when encoding these, but who knows...
 

Offline Bassman59

  • Super Contributor
  • ***
  • Posts: 2499
  • Country: us
  • Yes, I do this for a living
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #6 on: June 14, 2020, 04:45:25 am »
Awhile ago I did an S6 design at 100 MHz which used a lot of block RAMs, and I don't remember any particular timing challenges.
No fair!  ;)

Did you use them all combined into one large memory?

In a manner of speaking. The FPGA connects to eight ADC channels, each digitizing an output of an image sensor. The memory is a dual port, with a write side and a read side. (Not a "true dual port" where you can write to and read from each side independently.)

The write side looks like eight independent memories, as the output of each ADC has samples written to its associated buffer simultaneously. So say the line has 8192 pixels (and each pixel is 16 bits). Each buffer is 1024 pixels deep, so there is a ten-bit write address counter.

The read side looks like one big 8192-deep buffer, so it has a 13-bit address counter. The idea is that it de-interleaves the data "in place," so downstream functions just get lines with pixels already in order.

Actually there were two such buffers, accessed in a ping-pong fashion (writing to one while reading out of the other). This extends the size of the address counters by one bit, with the MSb being the X/Y select.

I ended up replicating the address counters for the write side, to limit the loading and the routing.

All of the memories were inferred, so no instantiation of elements from the Xilinx library.

Quote
Quote
Look at synthesis results. Did the synthesis do what you expect or did it do something completely wacky? Coding style is important. Remember that if you don't write code in a way the synthesizer expects (basically it's a template matching process) then you could get poor results. The other day I looked at the synthesis report for something and saw that it inferred a boatload of flip-flops instead of the block RAM I expected. That was from a mistake in the code that inferred the memory.
Thank you, that is a good point. I have not questioned the synthesis much, but focused on the place & route results mostly. I guess I will have to look into that processor core and understand it a bit less superficially, since that's where a lot of the action is. It is largely a largish state machine, and I have found that th synthesizer is quite clever when encoding these, but who knows...

It might just be that the processor core has long logic paths that can't be optimized. Of course if you wrote the core, or you can modify it, you can look for places where you can pipeline, but that might change how the core functions.
 

Offline OwO

  • Super Contributor
  • ***
  • Posts: 1248
  • Country: cn
  • RF Engineer.
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #7 on: June 14, 2020, 05:26:40 am »
Block ram without at least 2 pipeline registers is always going to be slow, and it doesn't get much better with new FPGA families. The block rams are far away and routing delay dominates. Can you post a screenshot of the failing path highlighted in planahead (the physical device view)?
Vivado optimizes better than ISE, but if you are on spartan 6 there isn't much you can do. I never had trouble getting >200MHz on the slowest speed grade spartan 6, but everything I did was easy to pipeline.
 

Offline ebastlerTopic starter

  • Super Contributor
  • ***
  • Posts: 7763
  • Country: de
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #8 on: June 14, 2020, 06:20:52 am »
Block ram without at least 2 pipeline registers is always going to be slow, and it doesn't get much better with new FPGA families. The block rams are far away and routing delay dominates.

Yes, that's exactly the bottleneck I face: The CPU softcore sits more or less in the center of the chip. (It needs a fraction of the logic cells only.) Since I need to use all 32 RAM blocks, that incudes the blocks at the far ends of the chip, and the route delays to those dominate the timing.

Quote
Vivado optimizes better than ISE, but if you are on spartan 6 there isn't much you can do.

Hence my hope that I could give ISE some "direction" via constraints, to get it to optimize somewhat better. But apparently there's not much one can do along those lines. Maybe I'll find some places for mild pipelining in the design...

It's annoying that Xilinx decided not to support Spartan-6 in Vivado. But I guess it makes sense from their perspective: They don't want Spartan-6 to be used for new designs and only keep selling it for use in existing devices. Then ISE is adequate, since all people are supposed to do with it is sustain their old designs.
 

Offline xlnx

  • Contributor
  • Posts: 20
  • Country: no
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #9 on: June 14, 2020, 10:12:36 am »
I suggest (from the synthesis results) you take an extra look at how the BRAMs are organized. If they are organized as deep and narrow (down to 1 bit wide each), then much less multiplexing of the datalines is required, which could hugely impact the combinatorial delays.

Even if you infer all your BRAMs, take a look at how CoreGen would build your 64kB BRAM, and try to do similarly when you infer it through your own code.
 

Offline asmi

  • Super Contributor
  • ***
  • Posts: 3330
  • Country: ca
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #10 on: June 14, 2020, 11:32:58 am »
1. Have you tried placing it into larger device just to see if maybe it's a routing congestion issue?
2. Do you have access to the source code for your CPU core? If so, adding one more pipeline stage is usually fairly trivial. Just watch out for feedback paths (signals which cross stage boundaries "the wrong way around").
3. Have you considering migrating to Spartan-7? They are significantly faster than S6, they are cheaper, and you can use Xilinx's Microblaze core for free if that's what you want.

Offline ebastlerTopic starter

  • Super Contributor
  • ***
  • Posts: 7763
  • Country: de
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #11 on: June 14, 2020, 01:31:00 pm »
I suggest (from the synthesis results) you take an extra look at how the BRAMs are organized. If they are organized as deep and narrow (down to 1 bit wide each), then much less multiplexing of the datalines is required, which could hugely impact the combinatorial delays.

I had that thought too, and tried explicitly declaring a set of eight 64k*1bit RAMs, and generating them via the IP Wizard. That did indeed win me 5 MHz initially, but again, the benefit has not been reproducible. Comes and goes with every new synthesis and place&route run...
 

Offline ebastlerTopic starter

  • Super Contributor
  • ***
  • Posts: 7763
  • Country: de
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #12 on: June 14, 2020, 01:45:25 pm »
1. Have you tried placing it into larger device just to see if maybe it's a routing congestion issue?

I have indeed tried to go from the LX9 to the LX16, and it brings a slight advantage. Not due to routing congestion with the LX9, it seems, but because the LX9 chip has a longish, rectangular aspect ratio, while the LX16 is more or less a square. That's how the "physical" view of the tools shows it, at least, and it seems to be real: In the LX16, the BRAM blocks at the far edges of the chip don't need to be used, and hence the maximum routes are shorter.

Quote
2. Do you have access to the source code for your CPU core? If so, adding one more pipeline stage is usually fairly trivial. Just watch out for feedback paths (signals which cross stage boundaries "the wrong way around").

Yes, it's open source. I had hoped to avoid digging into it, and more importantly, want to keep it cycle-correct when operating on the slow external bus. In principle that should be possible -- whenever an external bus access is needed, there's enough time to let the CPU catch up and empty its pipeline. Sounds a bit intricate and error-prone (on my end ;-) though...

Quote

3. Have you considering migrating to Spartan-7? They are significantly faster than S6, they are cheaper, and you can use Xilinx's Microblaze core for free if that's what you want.

Duh... no, I had not considered that yet. And I had not even realized that the S7 is lower cost than the S6; had always assumed it was the other way round by a fair margin. Better chip performance, modern toolchain, lower cost -- seems worth a look indeed, thank you!
 

Offline ebastlerTopic starter

  • Super Contributor
  • ***
  • Posts: 7763
  • Country: de
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #13 on: June 14, 2020, 02:47:38 pm »
Perhaps it is worth adding a few placement constraints to the design.

If only I knew where...  ::)

As mentioned in the original post, I did try placement constraints. Having looked at the timing report and the critical signal paths, I had noted that some of the CPU softcore's components had been placed pretty far away from the center, adding to the routing delays. So I restricted them to a pblock in a central area.

This helped somewhat, but only initially. With some minor changes in other corners of the design, the allowable clock rate deteriorated -- and I found that removing the placement constraint made things better again, presumably because the placer found better uses for the central area. I have yet to find a placement constraint that is so strategic in nature that it will lead to robust speed improvements.

Maybe the "solution" is to keep this kind of fiddling to the very end, when nothing is expected to change anymore in the design?
 

Offline asmi

  • Super Contributor
  • ***
  • Posts: 3330
  • Country: ca
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #14 on: June 14, 2020, 03:04:56 pm »
Duh... no, I had not considered that yet. And I had not even realized that the S7 is lower cost than the S6; had always assumed it was the other way round by a fair margin. Better chip performance, modern toolchain, lower cost -- seems worth a look indeed, thank you!
It's kind of a wash at the very low end (LX4 is cheaper than S7-6, but the latter have more resources), but everything else is in favor of S7 - S7-15 is about the same price as S6-9 (the price can go either way depending on the package, speed grade, but the former has 15k "gates" vs 9k in the latter), and as you go up the ladder the difference only grows.
Give it a try as it doesn't cost you anything (except for some of your time) - import your code into Vivado and run P&R just to see what happens. One thing you've got to keep in mind though is Vivado works differently than ISE in a sense that it's timing constraints-driven - meaning it will do it's best to close the timing, but not an iota more. The reason for that I suspect is that Vivado supports gigantic devices with insane amount of resources, so there is an exponential explosion of possible configurations and consequently it will take forever to try them all, so it settles on "good enough" solution that meets all constraints. This is why proper constraints are even more important for Vivado. But the downside of that approach is that the only way to see how far can you push your design is to progressively tighten timing constraints and seeing if P&R succeeds or not. But even then there are different strategies you can try by forcing P&R to try harder to find a solution. As you get close to the limit, you will see the P&R taking longer and longer as it needs to work harder to close timing.

Online NorthGuy

  • Super Contributor
  • ***
  • Posts: 3526
  • Country: ca
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #15 on: June 14, 2020, 05:47:37 pm »
I don't think routing delays prevent you from running at 100 MHz. 100 MHz means 10ns between clock edges. This is plenty for routing.

Look at the critical path and see what's in it. Do you really see only wires, or are there any LUTs on the way? How many layers of LUTs do you have?

BRAM blocks have optional output registers. Are they used in your design? If you use these, depending on your speed grade, it may save you several ns. Of course this will introduce one clock delay, so you may need to change the design, and few MHz may not be worth it.
 

Offline ebastlerTopic starter

  • Super Contributor
  • ***
  • Posts: 7763
  • Country: de
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #16 on: June 14, 2020, 05:55:06 pm »
I don't think routing delays prevent you from running at 100 MHz. 100 MHz means 10ns between clock edges. This is plenty for routing.

Look at the critical path and see what's in it. Do you really see only wires, or are there any LUTs on the way? How many layers of LUTs do you have?

Believe me, I have looked long and hard... Typically 60..70% of the critical path duration is due to routing delays.

Quote
BRAM blocks have optional output registers. Are they used in your design? If you use these, depending on your speed grade, it may save you several ns. Of course this will introduce one clock delay, so you may need to change the design, and few MHz may not be worth it.

Right, that's going back to the question of cycle-correct simulation of the processor -- a 65x02, by the way. I don't use the output registers at the moment, since the additional delay cycle would significantly change the processor's behavior. It typically wants those data right away, e.g. to calculate the address for the very next bus cycle.
 

Online NorthGuy

  • Super Contributor
  • ***
  • Posts: 3526
  • Country: ca
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #17 on: June 14, 2020, 06:37:40 pm »
Believe me, I have looked long and hard... Typically 60..70% of the critical path duration is due to routing delays.

That's not how I would look at it. When you add LUTs, you automatically add routing to and from LUTs. Routing delays are roughly determined by the number of switch boxes on the way. Any intermediary LUT will add at least two switching boxes to the route which might be a bigger delay than the LUT itself. If you have any combinatorial logic on the way, the solution is to minimize the number of combinatorial layers, which will also remove the associated routing.

For example, you want to drive from Los Angeles to San Francisco. If you decide to visit Las Vegas on the way, your trip will get longer and you will have to drive more miles. If you want to drive less, the solution is not to find better shorter roads, but to drop Las Vegas entirely.

Although you have the unregistered BRAM (3.5 ns delay for -1 speed grade), you still have plenty of time left for routing.
 

Offline ebastlerTopic starter

  • Super Contributor
  • ***
  • Posts: 7763
  • Country: de
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #18 on: June 14, 2020, 06:51:00 pm »
That's not how I would look at it. When you add LUTs, you automatically add routing to and from LUTs. Routing delays are roughly determined by the number of switch boxes on the way. Any intermediary LUT will add at least two switching boxes to the route which might be a bigger delay than the LUT itself. If you have any combinatorial logic on the way, the solution is to minimize the number of combinatorial layers, which will also remove the associated routing.

For example, you want to drive from Los Angeles to San Francisco. If you decide to visit Las Vegas on the way, your trip will get longer and you will have to drive more miles. If you want to drive less, the solution is not to find better shorter roads, but to drop Las Vegas entirely.

Although you have the unregistered BRAM (3.5 ns delay for -1 speed grade), you still have plenty of time left for routing.

Well, but having the LUTs (i.e. the logic for a working CPU) is the whole point of the design, isn't it? I can't really do away with them; just moving unmodified data from BRAM to BRAM wouldn't quite cut it. ;) 

I can't remove the LUTs. I could, in principle, introduce more pipelining stages, but am reluctant to do so since I don't want to mess with the original 6502 timing. So reducing the routing delays by optimal positioning of the RAMs and logic seems like an important piece of the optimization; and I had hoped that I could nudge the Xilinx tools a bit to make them find the optimum more reproducibly.

Anyway, I think I will simply settle for slightly less for now, and move on to more rewarding tasks. 96 MHz is also a nice clock rate since it aligns well with the 48 MHz USB clock...  8)
 

Online NorthGuy

  • Super Contributor
  • ***
  • Posts: 3526
  • Country: ca
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #19 on: June 14, 2020, 07:02:23 pm »
Well, but having the LUTs (i.e. the logic for a working CPU) is the whole point of the design, isn't it? I can't really do away with them; just moving unmodified data from BRAM to BRAM wouldn't quite cut it. ;) 

Sure, but you can do a better design which does the same with less combinatorial layers. Or you can move some of the combinatorial logic from the critical path to other stages of your pipeline. There are many things you can do. Of course, if you do not want to make any changes then nothing will change  ;)

You can improve timing by moving things around and doing manual placing and routing, but this is very tedious job. It would never occur to me to do this to gain 4 MHz.  Get a faster speed grade, or move to S-7.
 

Offline ebastlerTopic starter

  • Super Contributor
  • ***
  • Posts: 7763
  • Country: de
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #20 on: June 14, 2020, 07:18:34 pm »
You can improve timing by moving things around and doing manual placing and routing, but this is very tedious job. It would never occur to me to do this to gain 4 MHz.  Get a faster speed grade, or move to S-7.

Sure, the difference between 96 to 100 MHz would just be cosmetic in nature -- as I said in the OP, 100 is a three-digit number. :)

I am already using the -3 speed grade, so the XC6SLX25 (for slightly better BRAM placement) or Spartan-7 are the possible hardware upgrades. Later, maybe; for the moment I'll just swallow my pride and live with the two digits...
 

Online NorthGuy

  • Super Contributor
  • ***
  • Posts: 3526
  • Country: ca
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #21 on: June 14, 2020, 07:47:58 pm »
Sure, the difference between 96 to 100 MHz would just be cosmetic in nature -- as I said in the OP, 100 is a three-digit number. :)

I am already using the -3 speed grade, so the XC6SLX25 (for slightly better BRAM placement) or Spartan-7 are the possible hardware upgrades. Later, maybe; for the moment I'll just swallow my pride and live with the two digits...

In practical terms, if it closes at 96 MHz, it'll work at 100 MHz too. Just cheat it by saying your input clock is 24 MHz, but in reality feed it 25 MHz, or something like that. Chances that 4% speed increase kills it are slim to none.
 

Online SiliconWizard

  • Super Contributor
  • ***
  • Posts: 17795
  • Country: fr
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #22 on: June 14, 2020, 09:39:47 pm »
I don't think routing delays prevent you from running at 100 MHz. 100 MHz means 10ns between clock edges. This is plenty for routing.

Uh yeah. I have a design on a Spartan 6 that is not even that complex (but it's a CPU). I get a 110MHz max frequency, and among the longest paths, routing delay alone takes up to 7ns. My design uses 23% of the FPGA resources, so it's far from congested. I use almost all block RAM available though.
« Last Edit: June 14, 2020, 09:41:37 pm by SiliconWizard »
 

Online tom66

  • Super Contributor
  • ***
  • Posts: 8840
  • Country: gb
  • Professional HW / FPGA / Embedded Engr. & Hobbyist
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #23 on: June 14, 2020, 09:43:38 pm »
Agreed. I've run into timing issues at 125MHz on a Zynq 7014S with 30% utilisation.  Lots of AXI buses though.
 

Offline Harvs

  • Super Contributor
  • ***
  • Posts: 1210
  • Country: au
Re: Xilinx ISE - how to improve erratic timing optimization?
« Reply #24 on: June 14, 2020, 10:10:36 pm »
Duh... no, I had not considered that yet. And I had not even realized that the S7 is lower cost than the S6; had always assumed it was the other way round by a fair margin. Better chip performance, modern toolchain, lower cost -- seems worth a look indeed, thank you!
It's kind of a wash at the very low end (LX4 is cheaper than S7-6, but the latter have more resources), but everything else is in favor of S7 - S7-15 is about the same price as S6-9 (the price can go either way depending on the package, speed grade, but the former has 15k "gates" vs 9k in the latter), and as you go up the ladder the difference only grows.

I'm super interested to hear that, I'd been sticking with S6 and ISE in a VM just down to cost.
Current pricing on the S6 - 16k is <$6USD Qty 1.
https://lcsc.com/product-detail/CPLD-FPGA_XILINX_XC6SLX16-2FTG256C_XILINX-XC6SLX16-2FTG256C_C39313.html

Where can we get this sort of pricing on S7?
 


Share me

Digg  Facebook  SlashDot  Delicious  Technorati  Twitter  Google  Yahoo
Smf

 

-->