This is often used when peripheral can be pointed at resources which are in main memory, and that peripheral will pull those resources from main memory as needed. One obvious example is video card - you load models/textures/sprites/etc into RAM, point video card at them, an it does the rest on it's own.
There are more uses for bus mastering, I only provided a single example to get a taste of what's possible.
Yes bidirectional shared memory is certainly useful. Graphics cards make a lot of use for it for paging system ram into VRAM, loading data directly from SSD, talking between multiple GPUs..etc. Very useful when you have gigabytes of RAM and can't afford to be constantly copying gigabytes of data across the bridge.
But again a typical MCU is only going to have 10s or 100s of KB of RAM, so easier to keep any hot data being worked on inside FPGAs RAM and access it over the bus when needed. Or if you need it in the MCU RAM then just DMA it across in because the small amount of data that can fit in the MCU RAM will transfer in sub milisecond times.
Yes it's not hard when protocol is well-documented and hardware does actually work like documentation says it does. The problem is routing wide busses in not an easy task, which is why pretty much all modern high-speed interfaces are serial and not parallel. These also eat FPGA's IO pins like there is no tomorrow, which might force you to choose larger package than the one you would've chosen otherwise.
Yeah hence the use of SPI,QSPI,OSPI. Since there is SRAM chips for these buses means that some MCUs have native memory mapped support for these.
What's the point of connecting a 100 MHz MCU to FPGA if you can simply implement said MCU inside of FPGA via softcore which would run at similar frequency and have a simple and straightforward single chip solution instead of dealing with interconnects? I know in Vivado you get pretty much all peripherals you would find in an MCU available for free, and - even better - you can add exactly the peripherals you need and nothing more. Or you can add a lot of some peripherals than what you'd find in MCU if that's what you need. And such softcore would still have this glue-less interface to your custom peripherals, which can access main memory as well as other MMIO resources if configured to do so. And with Vivado you don't need to write a single line of HDL to build such a system from zero to a bitstream ready for programming (or firmware integration whatever the case may be).
For when you need to keep BOM cost down.
FPGAs are expensive chips as it is. So in order to justify one you already need a application that can't be solved using a reasonable amount of off the shelf single purpose chips or a MCU peripheral. So then once you know you need a FPGA you find that the FPGAs in the low 10s of dollar range are not very big. So stuffing your whole design including a CPU into the FPGA might not be viable. So the easiest step to reduce the logic inside is to buy an off the shelf MCU chip (since those are incredibly common) so you don't need a CPU inside the FPGA anymore. Then since a MCU also has peripherals you use those for as many tasks as they can do. This leaves you with only the funky non standard peripherals inside the FPGA so the whole FPGA can be dedicated to that.
Then since some of your machinery is inside the FPGA you need to talk to it somehow, hence you need a bus between the FPGA and MCU. Sometimes the traffic requirements are low (ie FPGA catapulting 1GB of streaming data and outputting a 1KB measurement result and some control commands back) so all you need is some pedestrian serial bus. In a FPGA implementing I2C or UART is a bit annoying so the choose is often SPI as it is just a simple shift register. But if you need more speed the easy step up is QSPI, then since some MCUs support memory mapped QSPI (due to existence of QSPI SRAM) it can easily be upgraded into a memory mapped bus to simplify drivers on the MCU side and reduce overhead while on the FPGA side there is not much extra complexity as you still get read/write commands with an address and data.
Hence by the end of that design process we invented a FPGA+MCU memory mapped bus solution.
This is the solution i used for when i needed to talk to high speed ADCs or do video compositing on MIPI DSI and LVDS video etc.. where i need to work with high speed data streams but i don't actually need any compute heavy processing done on it, yet i still need a CPU in there to set things up and orchestrate it all.