<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Audio | Romain Lopez</title>
    <link>https://romain-lopez.github.io/tags/audio/</link>
      <atom:link href="https://romain-lopez.github.io/tags/audio/index.xml" rel="self" type="application/rss+xml" />
    <description>Audio</description>
    <generator>Source Themes Academic (https://sourcethemes.com/academic/)</generator><language>en-us</language><copyright>Copyright © Romain Lopez 2026</copyright><lastBuildDate>Fri, 24 Apr 2026 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://romain-lopez.github.io/img/icon-192.png</url>
      <title>Audio</title>
      <link>https://romain-lopez.github.io/tags/audio/</link>
    </image>
    
    <item>
      <title>Wassersound: morphing sound by moving mass</title>
      <link>https://romain-lopez.github.io/post/wassersound/</link>
      <pubDate>Fri, 24 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://romain-lopez.github.io/post/wassersound/</guid>
      <description>&lt;!---
Placeholder for a hero image: a wide figure showing the two source spectrograms on the left and right, with the morphed spectrogram in between, annotated with the interpolation parameter λ running from 0 to 1.
--&gt;
&lt;p&gt;I&amp;rsquo;m teaching for the first time in ages this Fall! The class is a graduate seminar course on &lt;a href=&#34;https://romain-lopez.github.io/learning-to-match-distributions/&#34;&gt;distribution matching and morphing&lt;/a&gt;. While figuring out a few fun applications of computational optimal transport, I was sent back to a pandemic project (6 years ago!) with my friend Armand Bernardi. Armand had been long thinking about algorithmic and digital approaches to music composition. I was intrigued about a (then) &lt;a href=&#34;https://doi.org/10.1016/j.cell.2019.01.006&#34;&gt;recent paper&lt;/a&gt; of Geoffrey Schiebinger describing a quantitative theory of cellular development based on optimal transport. We explored how optimal transport could be used as a subroutine for algorithmic sound effect generation, playing around with existing work I refer to throughout the blog post.&lt;/p&gt;
&lt;p&gt;I loved working on this project so decided to talk about it here. On top of learning from Armand about algorithmic music generation, and having a fun few days hacking, this project ended up matching a pattern. Many real-world problems can be modeled with elegant mathematical formalism, but the real-world details are most often messy and the hardest to crack.&lt;/p&gt;
&lt;h2 id=&#34;wassersound&#34;&gt;Wassersound&lt;/h2&gt;
&lt;p&gt;A simple and natural place to start is a &lt;em&gt;glissando&lt;/em&gt;. Think of a violin dramatically &lt;a href=&#34;https://www.youtube.com/watch?v=vU2EciADd8o&#34;&gt;joining two notes&lt;/a&gt; together pulling the string in the right way &amp;ndash; got it? Optimal transport is a mathematical framework that can play the role of the violinist&amp;rsquo;s fingers, and continuously morph one sound into another. Specifically, each sound sample is spectrum over frequencies (e.g., the representation of the sound wave in Fourier space), and OT will tell you how to map frequency bins from one sound to another, directly enabling &lt;em&gt;morphing&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;For this article, I recorded two samples on the shakuhachi, a japanese bamboo flute I study (perhaps the topic of a later post!):&lt;/p&gt;
&lt;!--- AUDIO PLACEHOLDER: source A --&gt;
&lt;p&gt;&lt;strong&gt;Source A&lt;/strong&gt; — Tsu Re:&lt;/p&gt;
&lt;div style=&#34;text-align: center;&#34;&gt;
  &lt;audio controls src=&#34;https://romain-lopez.github.io/audio/wassersound/source_a.wav&#34;&gt;&lt;/audio&gt;
&lt;/div&gt;
&lt;!--- AUDIO PLACEHOLDER: source B --&gt;
&lt;p&gt;&lt;strong&gt;Source B&lt;/strong&gt; — Ro Chi Re Chi:&lt;/p&gt;
&lt;div style=&#34;text-align: center;&#34;&gt;
  &lt;audio controls src=&#34;https://romain-lopez.github.io/audio/wassersound/source_b.wav&#34;&gt;&lt;/audio&gt;
&lt;/div&gt;
&lt;h2 id=&#34;morphing-by-amplitude&#34;&gt;Morphing by amplitude&lt;/h2&gt;
&lt;p&gt;Linearly interpolating two raw audio waveforms — represented by an amplitude over time — just &lt;em&gt;crossfades&lt;/em&gt; them. Metaphorically, that&amp;rsquo;s what your cousin DJ-ing at a birthday party and transitioning between soundtracks would do. Let&amp;rsquo;s listen to it.&lt;/p&gt;
&lt;!--- AUDIO PLACEHOLDER: crossfade baseline --&gt;
&lt;div style=&#34;text-align: center;&#34;&gt;
  &lt;audio controls src=&#34;https://romain-lopez.github.io/audio/wassersound/crossfade.wav&#34;&gt;&lt;/audio&gt;
&lt;/div&gt;
&lt;p&gt;As expected, during the transition period, we get &lt;em&gt;superposition&lt;/em&gt; of the two tracks, and at the midpoint both notes are audible simultaneously. Therefore, you perceive a chord rather than a single pitch in transition. What would be musically more appealing is a signal that, at the midpoint, sounds like one note whose pitch lies between A&amp;rsquo;s and B&amp;rsquo;s. That is a statement about interpolating &lt;em&gt;distributions of spectral energy&lt;/em&gt;, not about interpolating waveforms.&lt;/p&gt;
&lt;h2 id=&#34;moving-into-fourier-space&#34;&gt;Moving into Fourier space&lt;/h2&gt;
&lt;p&gt;Our next move is therefore to go to the frequency domain. We use a &lt;a href=&#34;https://en.wikipedia.org/wiki/Short-time_Fourier_transform&#34;&gt;short-time Fourier transform&lt;/a&gt; (STFT): slide a 50 ms Hann window across each source in steps of half a window, and take an FFT at each position. The result is a complex spectrogram $S \in \mathbb{C}^{T \times N}$, with $T$ time-frames along the horizontal axis and $N$ positive-frequency bins along the vertical (we use a 50 ms window with 8× zero-padding, but the exact bin count doesn&amp;rsquo;t matter for what follows). Each column $S_{t.}$ is a snapshot of &lt;em&gt;which frequencies are present, at what amplitude and phase, during frame $t$&lt;/em&gt;.&lt;/p&gt;
&lt;!--- FIGURE PLACEHOLDER: 2-panel log-magnitude spectrograms of source A (top) and source B (bottom), with frequency on the y axis and time on the x axis. Same scale for both. --&gt;
&lt;p&gt;&lt;img src=&#34;https://romain-lopez.github.io/img/wassersound/source_spectrograms.png&#34; alt=&#34;Log-magnitude spectrograms of source A (top) and source B (bottom)&#34;&gt;&lt;/p&gt;
&lt;p&gt;Both sources are shakuhachi notes with roughly stationary timbre: you can see clean horizontal bands corresponding to a fundamental and its harmonics, fairly stable in time but at different frequencies for A and B. This is exactly the kind of signal where frame-by-frame magnitude spectra are meaningful objects, and where morphing them by moving mass along the frequency axis has a clean physical interpretation.&lt;/p&gt;
&lt;p&gt;We then turn each frame $S_{t.}$ of the audio sample into a &lt;em&gt;probability distribution over frequency&lt;/em&gt; $p_{t.}$:&lt;/p&gt;
&lt;p&gt;$$
p_{t,k} = \frac{|S_{tk}|}{\sum_{k&#39;} |S_{tk&#39;}|}.
$$&lt;/p&gt;
&lt;p&gt;At each of the $T$ time-frames, we now have $N$-dimensional probability vectors, $p_t^A$ and $p_t^B$ for each audio signal $A$ and $B$. The morphing problem reduces to: &lt;em&gt;for each $t$, interpolate between $p_t^A$ and $p_t^B$ using optimal transport.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&#34;optimal-transport-and-cost-functions&#34;&gt;Optimal transport and cost functions&lt;/h2&gt;
&lt;p&gt;Given two discrete distributions $\mu^0$ and $\mu^1$ on a shared ground space (here, our frequency bins) and a cost $C_{ij}$ for moving a unit of mass from bin $i$ to bin $j$, optimal transport finds the transport plan $\pi \in \mathbb{R}_+^{N \times N}$ that minimises&lt;/p&gt;
&lt;p&gt;$$
\langle C, \pi \rangle = \sum_{i, j} C_{ij}\pi_{ij},
$$&lt;/p&gt;
&lt;p&gt;subject to the positivity constraints for $\pi$, as well as the following marginal constraints&lt;/p&gt;
&lt;p&gt;$$\pi \cdot \mathbb{1} = \mu^0,$$&lt;/p&gt;
&lt;p&gt;and
$$ \pi^\top \mathbb{1}  = \mu^1.
$$
Those constraints force $\pi$ to be a valid &lt;em&gt;coupling&lt;/em&gt; of $\mu_0$ and $\mu_1$. We are especially interested in the minimiser $\pi^*$ which we call the &lt;em&gt;transport plan&lt;/em&gt;. This problem can be approximately solved quickly with an entropic regularisation term via the &lt;a href=&#34;https://arxiv.org/abs/1306.0895&#34;&gt;Sinkhorn algorithm&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The piece that matters for sound design is &lt;strong&gt;the cost matrix&lt;/strong&gt; $C$. $C_{ij}$ is our model of &lt;em&gt;&amp;ldquo;how far apart is frequency $i$ from frequency $j$?&amp;quot;&lt;/em&gt;, and the answer shapes the morph. We experimented with two choices.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Linear cost in frequency.&lt;/strong&gt; The simplest option: $C_{ij} = |i - j|$, normalised to $[0, 1]$. This just says that moving mass between nearby frequency bins is cheap and across the spectrum is expensive.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Harmonic cost.&lt;/strong&gt; A musically smarter choice, borrowed from &lt;a href=&#34;https://papers.nips.cc/paper/6479-optimal-spectral-transportation-with-application-to-music-transcription&#34;&gt;Flamary et al. (2016)&lt;/a&gt;. A note at fundamental $f$ and a note at $2f, 3f, \ldots$ are perceptually related — they are partials of the same harmonic series. So moving mass from $f_i$ to a harmonic multiple of $f_j$ should be almost as cheap as moving it to $f_j$ itself:&lt;/p&gt;
&lt;p&gt;$$
C_{ij}^{\text{harm}} = \min\left( (f_i - f_j)^2, ; \min_{q \in [2, Q]} (f_i - q f_j)^2 + \varepsilon q \right).
$$&lt;/p&gt;
&lt;p&gt;The barrier term $\varepsilon q$ keeps the procedure well-defined by penalising higher harmonics slightly (we use $\varepsilon = 10$, $Q = 5$). The resulting cost matrix has a striking geometry: a main diagonal, plus secondary &amp;ldquo;rails&amp;rdquo; at $j = i/2, i/3, \ldots$ along which transport is almost free.&lt;/p&gt;
&lt;!--- FIGURE PLACEHOLDER: 2-panel — linear cost matrix $C_{ij} = |i-j|$ on the left (banded diagonal gradient), harmonic cost on the right (main diagonal plus harmonic rails). Bin index on both axes. --&gt;
&lt;p&gt;&lt;img src=&#34;https://romain-lopez.github.io/img/wassersound/cost_matrices.png&#34; alt=&#34;Linear cost (left) and harmonic cost (right) matrices&#34;&gt;&lt;/p&gt;
&lt;p&gt;We will listen to both at the end of this blog post!&lt;/p&gt;
&lt;p&gt;The last piece of machinery we need is the &lt;em&gt;McCann interpolant&lt;/em&gt; $\mu^\lambda$ between $\mu^0$ and $\mu^1$ at parameter $\lambda \in ]0, 1[$ — a barycenter in Wasserstein space from $\mu^0$ at $\lambda = 0$ to $\mu^1$ at $\lambda = 1$. For discrete measures on a bin grid with optimal plan $\pi^\star$, it has a clean closed form:&lt;/p&gt;
&lt;p&gt;$$
\mu^\lambda_m = \sum_{i=1}^N\sum_{j=1}^N \pi^\star_{ij} \cdot \mathbb{1}_{m = \lfloor (1-\lambda) i + \lambda j \rceil},
$$
where $\lfloor \cdot \rceil$ denotes rounding to the nearest integer. This is one appealing reason to use OT for morphing: it provides a principled continuous path from one distribution to the other.&lt;/p&gt;
&lt;h2 id=&#34;spectrogram-morphing-with-ot&#34;&gt;Spectrogram Morphing with OT&lt;/h2&gt;
&lt;p&gt;To morph between sounds A and B, we compute the spectrogram $q \in \mathbb{R}^{T \times N}$ by putting the pieces together. For each time-frame $t$, we define $q_t$ as:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Compute the OT plan $\pi^{\star,(t)}$ between $p_t^A$ and $p_t^B$&lt;/li&gt;
&lt;li&gt;Form the McCann interpolant for $\pi^{\star,(t)}$ at parameter $\lambda_t = t/T$&lt;/li&gt;
&lt;li&gt;Record $q_t = \mu^{\lambda_t}$ defined above&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Look at the results:&lt;/p&gt;
&lt;!--- FIGURE PLACEHOLDER: 4-panel log-magnitude spectrograms — source A | morphed q (linear cost) | morphed q (harmonic cost) | source B. Same colour scale and frequency axis. --&gt;
&lt;p&gt;&lt;img src=&#34;https://romain-lopez.github.io/img/wassersound/morphed_spectrograms.png&#34; alt=&#34;Source A, morphed spectrograms (linear and harmonic cost), source B&#34;&gt;&lt;/p&gt;
&lt;p&gt;For the linear cost, the harmonic bands drift smoothly from $A$&amp;rsquo;s frequencies to $B$&amp;rsquo;s across the horizontal axis, and each individual frame looks like a plausible spectrum. For the harmonic cost, we see many crossing patterns that emerge from allowing the harmonics to match even at higher frequencies!&lt;/p&gt;
&lt;h2 id=&#34;reality-catches-up-----where-is-the-phase&#34;&gt;Reality catches up &amp;ndash;  &amp;ldquo;Where is the phase?&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;Alright, how does it sound? Well, the catch is that a magnitude spectrum is not enough to reconstruct an audio signal — we also need the phase.&lt;/p&gt;
&lt;p&gt;The cheapest (and wrong) thing one can do is to copy the phase $\arg S^A_{tk}$ from source $A$ and reconstruct the interpolation spectrogram $I \in \mathbb{C}^{T \times N}$:&lt;/p&gt;
&lt;p&gt;$$
I_{tk} = q_{t,k} e^{i \arg S^A_{tk}}.
$$&lt;/p&gt;
&lt;p&gt;Then take an inverse FFT of each frame, overlap-add, render to a WAV. For the linear OT map, the reconstructed signal sounds like this:&lt;/p&gt;
&lt;!--- AUDIO PLACEHOLDER: naive morph --&gt;
&lt;div style=&#34;text-align: center;&#34;&gt;
  &lt;audio controls src=&#34;https://romain-lopez.github.io/audio/wassersound/incoherent_linear.wav&#34;&gt;&lt;/audio&gt;
&lt;/div&gt;
&lt;p&gt;Hmm, not great right? Somehow, the pitch interpolated by OT does seem to move properly, but the signal is smeared and warbling. Across consecutive frames the phase is no longer evolving coherently with the spectrum it is attached to, and the ear is extremely sensitive to that mismatch.&lt;/p&gt;
&lt;p&gt;It was at this point we did a literature search with Armand, and discovered this &lt;a href=&#34;https://arxiv.org/abs/1906.06763&#34;&gt;really nice paper&lt;/a&gt;. They got this exact setup working with a few tricks based on phase accumulation, and have some &lt;a href=&#34;https://soundcloud.com/audio_transport&#34;&gt;demos on soundcloud&lt;/a&gt; worth listening. I would mention that horizontal incoherence is a hard problem. Their samples, and ours after reimplementation of their &lt;a href=&#34;https://github.com/sportdeath/audio_transport&#34;&gt;C++ code&lt;/a&gt; in python, are far from perfect.&lt;/p&gt;
&lt;h2 id=&#34;phase-reconstruction&#34;&gt;Phase Reconstruction&lt;/h2&gt;
&lt;p&gt;The way out begins with noticing that we are not actually trying to recover the phase of a &amp;ldquo;true&amp;rdquo; signal — there is no ground-truth waveform whose magnitude spectrogram is $q$. What we want is &lt;em&gt;some&lt;/em&gt; signal $x$ whose STFT magnitude is close to $q$ and that is consistent with overlap-add reconstruction. This is exactly the problem &lt;a href=&#34;https://ieeexplore.ieee.org/document/1164317&#34;&gt;Griffin and Lim (1984)&lt;/a&gt; formulated forty years ago:&lt;/p&gt;
&lt;p&gt;$$
x^\star = \arg\min_x \big| ,|\text{STFT}(x)| - q, \big|_F^2.
$$&lt;/p&gt;
&lt;p&gt;Their algorithm alternates two steps until convergence: starting from a random phase, take the inverse STFT to obtain a candidate signal, recompute its STFT, replace the magnitude by the target $q$ while keeping the recomputed phase, and repeat. The objective is monotonically non-increasing and converges to a local minimum.&lt;/p&gt;
&lt;p&gt;This is the right relaxation for our morph: instead of borrowing phase from a signal whose magnitude is no longer there, we solve a phase-consistency problem directly on the synthesised target. The same approach was recently used in this &lt;a href=&#34;https://arxiv.org/abs/2502.15430&#34;&gt;paper on OT for music&lt;/a&gt;. Listen to the final results!&lt;/p&gt;
&lt;!--- AUDIO PLACEHOLDER: griffin-lim morph --&gt;
&lt;p&gt;&lt;em&gt;Linear cost with Griffin–Lim phase reconstruction&lt;/em&gt; :&lt;/p&gt;
&lt;div style=&#34;text-align: center;&#34;&gt;
  &lt;audio controls src=&#34;https://romain-lopez.github.io/audio/wassersound/griffin_lim_linear.wav&#34;&gt;&lt;/audio&gt;
&lt;/div&gt;
&lt;p&gt;&lt;em&gt;Harmonic cost with Griffin–Lim phase reconstruction&lt;/em&gt;:&lt;/p&gt;
&lt;div style=&#34;text-align: center;&#34;&gt;
  &lt;audio controls src=&#34;https://romain-lopez.github.io/audio/wassersound/griffin_lim_harmonic.wav&#34;&gt;&lt;/audio&gt;
&lt;/div&gt;
&lt;h2 id=&#34;conclusion&#34;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;That&amp;rsquo;s it for today. Six years after, what I remember most about this fun project is how hard it was to get the infrastructure around it right (audio encoding and import, FFT, and of course phase reconstruction). It was frustrating to spend hours with Armand wondering how spectrograms that look so beautiful on the screen could sound so bad on speakers.&lt;/p&gt;
&lt;p&gt;This small project may be good allegory for a research project. The clean equations are usually the easy part; the work that decides whether anything functions is in the seams between them — the data, the encoding, the post-processing. Building solid tools and codebases, learning the details of data processing and relating them to the reality of experiments, and checking with simple examples whether an algorithm has the expected output: these are the skills that compound across years of research and let you prototype quickly. Now back to teaching!&lt;/p&gt;
&lt;h4 id=&#34;acknowledgments&#34;&gt;Acknowledgments&lt;/h4&gt;
&lt;p&gt;Thanks to &lt;strong&gt;Armand Bernardi&lt;/strong&gt; for the long afternoons in 2020 listening to roughly two hundred warbling reconstructions. The reassigned-STFT analysis is a Python port of Trevor Henderson and Justin Solomon&amp;rsquo;s &lt;a href=&#34;https://github.com/sportdeath/audio_transport&#34;&gt;audio_transport&lt;/a&gt;, itself an implementation of &lt;a href=&#34;https://ieeexplore.ieee.org/document/382397&#34;&gt;Auger and Flandrin&amp;rsquo;s reassignment method&lt;/a&gt;. The harmonic cost is from &lt;a href=&#34;https://papers.nips.cc/paper/6479-optimal-spectral-transportation-with-application-to-music-transcription&#34;&gt;Flamary et al. (2016)&lt;/a&gt;; OT solvers are from &lt;a href=&#34;https://pythonot.github.io/&#34;&gt;POT&lt;/a&gt;; phase reconstruction via &lt;a href=&#34;https://ieeexplore.ieee.org/document/1164317&#34;&gt;Griffin &amp;amp; Lim (1984)&lt;/a&gt; as implemented in &lt;a href=&#34;https://librosa.org/&#34;&gt;librosa&lt;/a&gt;.&lt;/p&gt;
</description>
    </item>
    
  </channel>
</rss>
