Awhile ago I did an S6 design at 100 MHz which used a lot of block RAMs, and I don't remember any particular timing challenges.
No fair! 
Did you use them all combined into one large memory?
In a manner of speaking. The FPGA connects to eight ADC channels, each digitizing an output of an image sensor. The memory is a dual port, with a write side and a read side. (Not a "true dual port" where you can write to and read from each side independently.)
The write side looks like eight independent memories, as the output of each ADC has samples written to its associated buffer simultaneously. So say the line has 8192 pixels (and each pixel is 16 bits). Each buffer is 1024 pixels deep, so there is a ten-bit write address counter.
The read side looks like one big 8192-deep buffer, so it has a 13-bit address counter. The idea is that it de-interleaves the data "in place," so downstream functions just get lines with pixels already in order.
Actually there were two such buffers, accessed in a ping-pong fashion (writing to one while reading out of the other). This extends the size of the address counters by one bit, with the MSb being the X/Y select.
I ended up replicating the address counters for the write side, to limit the loading and the routing.
All of the memories were inferred, so no instantiation of elements from the Xilinx library.
Look at synthesis results. Did the synthesis do what you expect or did it do something completely wacky? Coding style is important. Remember that if you don't write code in a way the synthesizer expects (basically it's a template matching process) then you could get poor results. The other day I looked at the synthesis report for something and saw that it inferred a boatload of flip-flops instead of the block RAM I expected. That was from a mistake in the code that inferred the memory.
Thank you, that is a good point. I have not questioned the synthesis much, but focused on the place & route results mostly. I guess I will have to look into that processor core and understand it a bit less superficially, since that's where a lot of the action is. It is largely a largish state machine, and I have found that th synthesizer is quite clever when encoding these, but who knows...
It might just be that the processor core has long logic paths that can't be optimized. Of course if you wrote the core, or you can modify it, you can look for places where you can pipeline, but that might change how the core functions.