Ardour 9.7 w/96k Sample Rate with JACK/Pipewire

Mike, even your ACM70SA sounds different at 96kHz, even without doing compression.

I suppose that’s entirely possible - the other general rule is that there are no rules really, often its 'whatever sounds best to you :slight_smile:

1 Like

Have you done a NULL test? or looked at the spectrum if there is anything meaningful above 24kHz?

Also keep in mind that many plugins upsample internally where it makes sense to avoid aliasing, this way you don’t have to run the session at a higher sample-rate. Some also have dedicated oversampling controls.

Amen!

Here it is:

Even your darc does a somewhat different job on the transients.
And we can argue if that is a meaningful difference or not.
In the case of the ACM70SA it is quite obvious.

Yes, the attack and release time constants are only approximated for performance reasons, they are not exact, but close enough.

…and most probably you ignored Ardour’s cpu frequency scaling message/warning at the start of the session as i also did some time ago :slight_smile: .

I suspected this problem was about that (i had the same issue on several machines), but i’ve kept my mouth shut since there’s so many much smarter people in this conversation that me. It’s amazing how many times i see people go over the hills and back just to get some water while the well is right next to them :slight_smile: . This is (from my perspective) the most common high dsp load problem with new, unprepared setups that always gets mentioned last. I mean - 12th gen i5, 32gb ram, 3070 :slight_smile: :slight_smile: :slight_smile: - i’ve edited 4K videos with heavy image processing with much much less than that.

Ah, never underestimate yourself, or overestimate either as I get in the habit of too often. I don’t think I’ve ever seen the warning on frequency scaling your mentioning, but probably just glossed over it (who reads that stuff anyway :slight_smile: ), and I don’t think I would have attached it to actual DSP performance, but that was certainly the case.

The light bulb went on for me a couple times in this thread and I definitely learned some stuff. It’s stuff I knew but just couldn’t pull it into context of the problem I was working on, and I didn’t fully understand how my CPU and OS were managing its frequency until I sat and watched its performance, and then the lights started going on. So don’t hesitate to share. Never know when the right word at the right time will break through. Thanks!

2 Likes

So we arrived to what I said: 96k is not just about things above 20kHz.
Wonders, mysteries and magic below 20kHz.
Sometimes here, sometimes there.

FWIW, I now usually run 96k. This is in order to reduce audible-to-me aliasing artifacts from saturation & compression plugins that don’t properly support oversampling, and to have lower latency, even with plugins that do have decent oversampling options. I typically end up with near-zero latency from plugins, except the limiter. I filter out ultrasonics between non-linear plugins (using Airwindows hypersonic.)

@MarkR1 Have you checked whether anything illuminating is reported in Ardour’s log window? I sometimes see the “ambiguous latency” warnings if there’s a webcam mic or similar plugged in, and ALSA or Pipewire is set to use that as an input. The culprit is usually obvious from the log messages. I vaguely recall something similar also happening with on-board audio (as well as other issues if on-board audio is set to “pro” profile).

BTW I’ve sometimes been obliged to track and mix parts of a 96k session on a 48k interface. That’s probably far from ideal, but it seems to work with no obvious artifacts (at least for me). I do my initial tracking with FLAC, which helps keeps the file sizes under control.

Argh, I replied to the wrong person, deleted it, and tried to reply correctly, and it blocks me cause it’s to close the last post. I really don’t like this forum and the limitations in it, but suppose it’s necessary anymore. Maybe this paragraph will do it… and it did.

@foxcj Hey, I have not seen any ambiguous latency messages lately since getting my system squared away. I don’t have a webcam, but could have been my headset when things were getting overwhelmed. The messages were actually ambiguous for me and didn’t say what was causing it. Things are operating a lot better now and I can sample at 96k with a 1024 buffer size and still stay below 10% DSP load, but that will probably start climbing as I can now finally quit fighting my system and get to music making, and will probably start using more plugins as I get towards final mixing… I hope.

BTW, for anyone who may wondering how to do some of this, here is the script I came up with to set the CPU in performance mode ready to respond to the buffer fill and set it back when done for regular day-to-day use:

powerprofilesctl set performance
ardour9
powerprofilesctl set balanced

Pretty simple. Since that is using a system daemon, it should work safely on just about any CPU and Debian based system, but I’ve only tested it on Kubuntu 24.04. It does the same thing as setting the power profile in the system settings power management section. I setup an icon in the menu with “bash” as the command and the full path and filename to the script file to run it with one click.

That doesn’t disable frequency scaling completely but sets your base frequency at the max CPU frequency so there’s no need for the governor to scale it higher during operation. It limits scaling pretty well, but I have noticed it downscale occasionally.

If you want to mess with completely disabling frequency scaling and locking the CPU to one frequency, I have scripts for that but don’t want to just dump them out as you might be able to hose your system performance up with them if you’re not careful and count your zero’s correctly, but if anyone’s interested, post and I’ll put them up here. They are a little more complex to use than the three liner above.

@MarkR1 For me the biggest improvement in stability at 96k was swapping my studio machine for one without that long-standing Intel USB3 bug (which made even 48k a bit iffy!), followed by disabling hyperthreading. I was already setting performance scheduling, and configuring real-time IRQ priorities. Disabling HT allows me to record reliably at 96k and a 512 buffer even on relatively weedy hardware, with a fairly low DSP load and no problematic XRUN-generating spikes. But if I use Pipewire and leave my webcam plugged in, then all hope is lost (but that’s probably a config issue on my part…)

In reality, I guess most listeners probably wouldn’t notice if I recorded most things with14 bits at 32k; done with care, that could be of better technical quality than a lot of old indie records, but in digital it’d be a pain managing gain staging, and any audible artifacts of non-linearities.

BTW, it’s quite interesting to crank Mixbus’ compressor [edit: “drive”] and push a sine sweep through it, comparing what you can actually hear at 48k vs 96k. To my ears, one of those has more birdies than a local dawn chorus… of course whether that really matters depends on what you’re doing, and what you’re expecting.

If you’re doing 48k vs 96k tests you should do it blindly and with matched volumes.
Confirmation bias is a big factor when it comes to “this audio sounds better that that”.

I mean, if you want to run everything at 96k it’s your prerogative but I’m pretty sure all you’re doing is wasting disk space.
As Robin said “Unless you make music for rats or dolphins there are only disadvantages of using sample-rates > 48kHz

And the obligatory XiphMont link, when it comes to 96k music :
https://people.xiph.org/~xiphmont/demo/neil-young.html

2 Likes

Yes, but what was described in the previous post is aliased distortion products. That won’t be subtle, it will be extra tones moving down in frequency as the sine sweep moves up in frequency. That should even show up on a spectrogram, so it shouldn’t be something on the edge of audibility that you would need carefully matched levels to be sure you are hearing.

I would never claim a 96k recording sounds better than a 48k recording (nor that a 24 bit one necessarily sounds better than one with 16 bit audio), but most people can hear the aliasing artifacts that arises from some non-linear plugins under some circumstances. Good plugins address this by oversampling to 96 or 192k (or beyond). But this adds latency, and load. Some people claim to be able to hear oversampling artifacts with some less-well implemented plugins due to pre-ring or phase-shifts (but such artifacts are not audible to my ill-trained ears).

Overall my latency is much lower running the whole session at 96k, as all the ‘oversampling’ is done ahead of time, and for the specific plugins that I use, my overall CPU load also turns out to be lower than when running at 48k and oversampling within plugins. The aliasing suppression at 96k with ultrasonic filtering is OK (but not perfect). For my use case, 96k FLACs turn out to be much smaller than 48k WAVs.

There is some theoretical and experimental work on low-latency techniques that suppress the generation of alias artifacts from non-linear processing without oversampling, but I don’t know of any decent plugins that do this sort of thing. (The ones I have tried are a bit ‘meh’, in my opinion.) For ultra low-latency tracking, I could use outboard gear, but that means more fragile hardware to cart around, and longer setup times. Or I could track at 48k, and then remix with oversampling, which is what I used to do, but for me and my workflow, that uses up valuable time, and rules out some otherwise nice-sounding plugins.

BTW, for the aliasing demonstration, I should have said try the Mixbus 12 drive in a 48 or 44.1k session, rather than the compressor. There are clearly audible birdies running up and down the spectrum given a sine sweep (which are also clearly visible on a spectrum display). These can easily peak above -50dBFS from a -6dBFS sweep, even without cranking the drive too hard. Running at 96k, these harmonics mostly just disappear into the ultrasonic part of spectrum, where they should be, and where they can be nicely trimmed off with an Airwindows hypersonic filter, or similar, before they start to trouble the local bat population, or any subsequent non-linear processing. Of course it’s nicer to do the comparison by attempting a null test, then you don’t have to listen to that awful tinnitus-inducing sweep!

One reason to go to a higher sampling rate is that the hardware anti-aliasing filters in place are rarely perfect. They are supposed to be flat to within 3db across the audio spectrum. They should start their roll off around 22.1khz. At that point, the signal has reached its 3db down point, but where it reaches the next 3db down point will depend on the quality of the filter. Cheaper filters will roll off at less than acceptable rates and leave the possibility for aliased frequencies to bleed through. And the higher quality of the filter, the more cost associated with it. You really have to look at the response curve of the input filter to know.

That’s one of the main reasons they started sampling at 48khz. 44.1khz was fine for final production, but for sampling the filter has be a nearly perfect square dropoff. Moving up to 48khz gives you more headroom, but you still need a pretty high pole low pass or band pass filter to ensure nothing bleeds through. Going to 96khz gives you plenty of headroom to eliminate all aliasing.

So it comes down to your interface. Some of them may have great filters and never alias at any sampling frequency, others may cut corners here to save a few bucks since the interface supports higher sampling frequencies. Seems the bottom line is use what’s best for your system.

I would agree though that it’s better to get enough samples to work with because there are some effects, like pitch shifting and time stretching, that can benefit from having more samples to work with. It will help minimize any artifacts from the processing. They often leave the sound sounding a bit more mechanical than before. To me at least. I have yet to really work with them at 96k to see if it’s actually any better, but as I said many posts ago, I’d rather have the extra samples and not need them, than need them and have to start over to get them.

If you’re able to hear that your audio hardware is seriously defective.
Unless it’s in the 96k case and your amp or speaker just isn’t able to handle the ultrasonic audio, as explained in the XiphMont post.

And just because we can see something in a spectrogram doesn’t mean that we can hear it.
We can see a 90 kHz tone in a spectrogram just fine, but we sure as hell can’t hear it.

Well, I’m pretty sure a 48k FLAC would be even smaller that the 96k one.
But it’s your disk space so…

Have you done a blind test to determine that you actually are able hear that?
I remember doing an ABX test on a “worst case”, and properly compressed, mp3 file and found out that somewhere above 140 kbps I was just guessing which was the original wav and which was the mp3.
And just because you can see a spectrum representation of something doesn’t mean that you can actually hear it, if it’s above like 15kHz or something (I presume that you’re at least in your 30’s or 40’s, meaning that you’ve probably lost the “up to 20kHz” hearing).

And I would argue that about 48 of those 96 kHz are completely unnecessary anyway, unless you have some really crappy audio hardware.
Or if you’re making music for dogs or bats.

But again; if someone’s OK with wasting disk space I won’t argue with their artistic choices.
I’ll just note that the Beatles were able to make some great hits using the equivalence of something like 10 bits/32kHz, in today’s rates.

If Monty’s Xiph Video doesn’t convince someone they can’t be convinced :wink: And with most productions destined for streaming…

1 Like

The original choice of 44100Hz as a sample rate is related to the fact that the only device capable of storing meaningful amounts of digital audio data when CD was invented (or technically prior to the invention of CD) was a VCR transport.
To enable reuse with minimal modification, this ran at the same speed as video, and used much of the same circuitry. 44.1 kHz was the highest usable rate compatible with both PAL and NTSC video with 3 samples per video line per audio channel.

e.g.

PAL:
294 Active lines per field
50 Fields per second
3 Samples per line
294 × 50 × 3 = 44,100 Hz

NTSC:
245 Active lines per field
60 Fields per second
3 Samples per line
245 × 60 × 3 = 44,100 Hz

2 Likes