RGMII for gigabit ehternet runs on a 125MHz clock, but uses double data rate technology, so on each clock edge the data has to be ready. This means data clocked every 4ns. Both the clock and data lines normally switch at the same time, but the RTL8211 allows for delaying the data lines of both the RX and TX side, reducing the need of delaying the clock signals. Based on this a design should still work, even if there is a trace length difference of an inch, or even more.
Yes, RTL8211 provides configurable TXDLY and RXDLY options that enable or disable a fixed 2 ns clock delay relative to the data. They don't provide fine-grained phase adjustment, but rather implement the clock-to-data offset required by the RGMII specification. This allows the required timing relationship to be generated inside the PHY instead of having to create these delays in the FPGA.
The RTL8211 datasheet guarantees at least 0.8 ns of hold time. Typical values may be larger, but timing closure should always be based on the guaranteed minimum values rather than typical ones. A fixed 2 ns delay option can be used when the overall timing already falls within the available margin, but it cannot compensate for arbitrary clock/data phase errors. It is intended to provide the fixed clock-to-data offset required by the RGMII specification, not to replace fine-grained timing adjustment.
I can agree, that from a pure propagation-delay perspective, 0.8 ns certainly looks like a fairly comfortable margin. At roughly 150-180 ps/inch, it corresponds to several inches of trace length difference. I assume that's the point you're making.
However, that estimate only considers propagation delay. In a real design, the sampling margin is also affected by signal integrity: trace bends, vias, impedance changes, parasitic capacitance and inductance, reflections, crosstalk and jitter. Those effects change not only the propagation delay, but also the edge shape and the instant at which the FPGA input actually crosses its logic threshold. As a result, the effective sampling margin can differ from what a simple propagation-delay estimate would suggest.
Returning to the original topic, the topic starter concern was about FPGA interfacing in general. ADC/DAC interfaces typically don't provide the configurable clock/data delay options at all.
In one of my prototypes (parallel interface clocked at 100 MHz), I initially connected the ADC to the FPGA using a short (~10-12 cm) ribbon cable. I remember that even relatively small changes to the clock wire affected sampling stability. Making the clock wire slightly longer/shorter (± 1-2 cm) or slightly changing its position/geometry while keeping approximately the same length was enough to shift the sampling point so that the FPGA started capturing the bus during transitions, producing obvious noise.
I could observe this effect in real time by gently bending the clock wire with my fingers. As soon as I bent the wire beyond a certain point, the captured test signal immediately became distorted and noisy. Straightening the wire restored stable operation. That's one of the reasons I've become quite cautious about wire lengths and routing.
Therefore, I'm skeptical of generalizing the conclusion that "an inch, or even more difference should work" to high-speed parallel interfaces. In practice, I have observed that even relatively small changes in clock routing at 100 MHz can have a noticeable impact on sampling stability in real hardware.