Author Topic: Reconstructing voices from spectrograms  (Read 1258 times)

0 Members and 1 Guest are viewing this topic.

Offline BudTopic starter

  • Super Contributor
  • ***
  • Posts: 7932
  • Country: ca
Reconstructing voices from spectrograms
« on: May 23, 2026, 08:53:32 pm »
NTSB has limited public access to their records that have dead pilot's voice spectrograms

https://arstechnica.com/ai/2026/05/ai-users-re-create-dead-pilots-voices-from-crash-investigation-docs

Apparently people (+AI) reconstruct crashed air planes cocpit audio using the spectrograms available from NTSB  - seems NTSB is not allowed to release actual audio.
I have never heard one reconstruction, but find this technically intriguing (you moral morons pls stay away from this thread, it is purely technical conversation).

Anyone seen this type of work done before?
« Last Edit: May 23, 2026, 09:01:01 pm by Bud »
Facebook-free life and Rigol-free shack.
 

Online tom66

  • Super Contributor
  • ***
  • Posts: 8843
  • Country: gb
  • Professional HW / FPGA / Embedded Engr. & Hobbyist
Re: Reconstructing voices from spectrograms
« Reply #1 on: May 23, 2026, 09:04:58 pm »
There's literally a video on YouTube for flight 2976 (the UPS flight that crashed after it literally lost an engine) that very convincingly reproduces the audio from the spectrograms.  Out of respect for the pilots I won't link it here, but you can find it if you really want.

The problem for the NTSB is most of these documents will be on Archive.org and have been reproduced elsewhere, as a US government agency everything they publish is public domain so they have no control over it once it leaves their server.
 

Offline I wanted a rude username

  • Frequent Contributor
  • **
  • Posts: 687
  • Country: au
  • ... but this username is also acceptable.
Re: Reconstructing voices from spectrograms
« Reply #2 on: May 23, 2026, 11:14:32 pm »
That's likely the very same one Bud's referring to: u/AlexandraMaryWindsor's post on Reddit, UPS flight 2976 CVR Spectrogram Reconstruction.

Supposedly it didn't require AI, as it's just performing the reverse transformation (back into the time domain). The possibility for this had been known for some time. The main constraint was the resolution of the image.
 

Offline pickle9000

  • Super Contributor
  • ***
  • Posts: 2441
  • Country: ca
Re: Reconstructing voices from spectrograms
« Reply #3 on: May 24, 2026, 02:01:53 am »
Scott Manley was apparently the one who started it. I hope he doesn't get in trouble, he has a great channel.

Link
 

Online nctnico

  • Super Contributor
  • ***
  • Posts: 30215
  • Country: nl
    • NCT Developments
Re: Reconstructing voices from spectrograms
« Reply #4 on: May 24, 2026, 01:15:21 pm »
NTSB has limited public access to their records that have dead pilot's voice spectrograms

https://arstechnica.com/ai/2026/05/ai-users-re-create-dead-pilots-voices-from-crash-investigation-docs

Apparently people (+AI) reconstruct crashed air planes cocpit audio using the spectrograms available from NTSB  - seems NTSB is not allowed to release actual audio.
I have never heard one reconstruction, but find this technically intriguing (you moral morons pls stay away from this thread, it is purely technical conversation).

Anyone seen this type of work done before?
Very possible. It is actually how audio stretching and pitch changing are based on. Calculate FFT (spectrum) over a short snippet, transpose and reconstruct.
There are small lies, big lies and then there is what is on the screen of your oscilloscope.
 

Online gf

  • Super Contributor
  • ***
  • Posts: 1831
  • Country: de
Re: Reconstructing voices from spectrograms
« Reply #5 on: May 24, 2026, 02:17:46 pm »
Very possible. It is actually how audio stretching and pitch changing are based on. Calculate FFT (spectrum) over a short snippet, transpose and reconstruct.

If you have the full, complex STFT (Magnitude AND Phase), perfect reconstruction is straightforward and trivial if COLA condition is met. Even if just NOLA condition is met you can divide out the resulting ripples. However, the challenge begins when you only have a spectrogram (which is just a visual image / matrix of the magnitude), because the phase information is completely gone.
 

Offline u666sa

  • Frequent Contributor
  • **
  • Posts: 893
  • Country: us
  • Miami, FL
    • Codernov Electronics Repair
Re: Reconstructing voices from spectrograms
« Reply #6 on: May 24, 2026, 03:01:58 pm »
There's literally a video on YouTube for flight 2976 (the UPS flight that crashed after it literally lost an engine) that very convincingly reproduces the audio from the spectrograms.  Out of respect for the pilots I won't link it here, but you can find it if you really want.
Voice confirms they were composed and tried to take plane away from populated areas.
 

Offline u666sa

  • Frequent Contributor
  • **
  • Posts: 893
  • Country: us
  • Miami, FL
    • Codernov Electronics Repair
Re: Reconstructing voices from spectrograms
« Reply #7 on: May 24, 2026, 03:04:35 pm »
It's a common practice among pilots to take doomed plane away from populated areas to minimize casualties.
 

Online SiliconWizard

  • Super Contributor
  • ***
  • Posts: 17798
  • Country: fr
Re: Reconstructing voices from spectrograms
« Reply #8 on: May 24, 2026, 04:29:40 pm »
Very possible. It is actually how audio stretching and pitch changing are based on. Calculate FFT (spectrum) over a short snippet, transpose and reconstruct.

If you have the full, complex STFT (Magnitude AND Phase), perfect reconstruction is straightforward and trivial if COLA condition is met. Even if just NOLA condition is met you can divide out the resulting ripples. However, the challenge begins when you only have a spectrogram (which is just a visual image / matrix of the magnitude), because the phase information is completely gone.

Yes, now getting anything that sounds like the original audio is almost a lost cause, from just a spectrogram, but getting something understandable is doable. The general way to do this is the "vocoder" approach. You can use inverse FFT, you can also just generate a pure sine or rather bandwidth-limited noise for each frequency band. The output won't sound nice but speech can usually be understood. Recognizing individual voices though is very hard, if this is the intent, but following a conversation is possible. The fact they only release spectograms is possible exactly for that reason: making individual identification very difficult if impossible, to protect the rights of the pilots.

 
The following users thanked this post: nctnico


Share me

Digg  Facebook  SlashDot  Delicious  Technorati  Twitter  Google  Yahoo
Smf

 

-->