Opus Audio Codec: SILK/CELT Hybrid, Low Latency for WebRTC, and FFmpeg Encoding

Key takeaways

IETF standard Opus: voice/music modes, SILK/CELT hybrid, low-latency real-time communication, and FFmpeg encoding. Master the next-generation royalty-free audio codec.

Introduction

Opus is an audio codec standardized as IETF RFC 6716 (2012), designed to handle both speech (the SILK layer) and music (the CELT layer) in a single bitstream. It covers roughly 6 to 510 kbit/s, frame lengths from 2.5 ms to 60 ms, and bandwidths from narrowband telephone audio up to full-band 20 kHz, and it can switch between all of these from one packet to the next without renegotiation. That flexibility is the real reason it won in real-time communication: a video call does not know in advance whether the next second will contain speech, music, silence or a burst of packet loss, and Opus can adapt to each without the application doing anything.

It is also published under a BSD license with royalty-free patent grants from its main contributors, which is why browsers, game engines and chat apps could ship it without per-unit fees. This guide explains how the codec is structured, what the knobs in FFmpeg’s libopus encoder actually change, where Opus is the right choice and where AAC or MP3 still make more sense.


Codec Overview

History and Development Background

Opus merged two existing codecs. SILK came from Skype, which had built it for VoIP over unreliable consumer internet connections; it is a linear-prediction speech codec that is very efficient on voice and poor on music. CELT, developed at Xiph.org, is a low-delay transform codec intended for music and interactive audio. The IETF codec working group combined them into one specification, so an application no longer had to choose between Speex or SILK for voice and Vorbis or AAC for music. After standardization in 2012 it became the mandatory-to-implement audio codec for WebRTC, and that single decision put it into every major browser.

Development has continued since. libopus 1.5 (2024) added machine-learning based features such as deep packet loss concealment and DRED (Deep REDundancy), which lets a sender embed a low-rate description of up to about a second of past audio so a receiver can reconstruct longer losses. These are optional and require both ends to support them, but they show where the codec is still improving: robustness on bad networks rather than raw compression.

Technical Features

ItemDescription
Compression MethodSpeech: linear prediction (SILK layer). Music and high frequencies: MDCT transform coding (CELT layer). Hybrid mode uses both at once
Sample RateInternal rates of 8, 12, 16, 24 and 48 kHz; Ogg Opus always decodes at 48 kHz
BitrateAbout 6–510 kbit/s; speech is usable from roughly 10–16 kbit/s, transparent stereo music typically needs 128 kbit/s or more
Frame Length2.5, 5, 10, 20, 40 or 60 ms per frame (20 ms is the usual default)
Algorithmic delay26.5 ms at the default 20 ms frame; can go down to about 5 ms with 2.5 ms CELT-only frames

Main Modes: SILK, Hybrid and CELT

Opus has three operating modes, and the encoder picks among them per packet based on bitrate, the requested bandwidth and the signal:

  • SILK-only: used for narrowband to wideband speech at low bitrates. Linear prediction models the vocal tract, so it spends bits efficiently on voice but handles music poorly.
  • Hybrid: SILK codes the band below 8 kHz and CELT codes the band above it. This is what lets speech at around 24–32 kbit/s sound “full-band” rather than telephone-like.
  • CELT-only: an MDCT transform codec used for music, for higher bitrates, and for the lowest-delay configurations (SILK does not support frames shorter than 10 ms).

In real-time communication the codec is only one part of perceived quality; jitter buffering, packet loss concealment and echo cancellation in the surrounding stack matter just as much.


Compression Principles

Psychoacoustic Model

The CELT layer is a perceptual transform codec: it relies on auditory masking to decide which spectral detail can be coarsely quantized. Its distinctive trick is that it explicitly preserves the energy of each band and codes only the shape of the spectrum within the band. When bits run out, the band still carries the right loudness, which is why low-bitrate CELT tends to sound slightly rough rather than muffled, a different failure mode from MP3’s “underwater” artifacts. The SILK layer instead models how speech is produced (pitch and the resonances of the vocal tract), which is far more efficient for voice and inappropriate for instruments.

MDCT (Modified Discrete Cosine Transform)

CELT uses the MDCT like AAC and Vorbis, but with much shorter frames and a low-overlap window. The short frames are what give Opus its low delay; the cost is lower frequency resolution, which CELT compensates for with techniques such as pitch prefiltering. This is why Opus at low delay is competitive on voice-over-music content but a long-frame codec like AAC can still be slightly more efficient for stored music at the same bitrate.

Bitrate Allocation Strategy

In the common VBR mode the encoder varies packet size with content complexity, analyses the input to decide between speech-oriented and music-oriented coding, and chooses the audio bandwidth for the available bitrate. As a user you mostly steer it indirectly, through the target bitrate, the application hint, the frame duration and the channel count; forcing a mode is rarely a good idea.

Processing Flow (Conceptual)

flowchart TB
  IN["PCM Input"]
  DET["Voice/Music Path Selection"]
  SILK["SILK Family Processing"]
  CELT["CELT (MDCT) Processing"]
  MIX["Bitstream Packing"]
  OUT["Opus Frame"]
  IN --> DET
  DET --> SILK
  DET --> CELT
  SILK --> MIX
  CELT --> MIX
  MIX --> OUT

Practical Encoding

FFmpeg Examples (Various Bitrates & Quality)

Ogg Opus, voice/podcast mono 32 kbps

ffmpeg -i input.wav -c:a libopus -b:a 32k -ac 1 -ar 48000 voice.opus

Music stereo 128 kbps (typical archive/distribution starting point)

ffmpeg -i input.wav -c:a libopus -b:a 128k -ac 2 -ar 48000 music.opus

Higher quality music 160~192 kbps

ffmpeg -i input.wav -c:a libopus -b:a 192k -ac 2 -ar 48000 music_hq.opus

WebM container (with video pipeline)

ffmpeg -i input.mkv -c:v copy -c:a libopus -b:a 96k -ac 2 output.webm

Use -c:a libopus, not FFmpeg’s native -c:a opus encoder. The native one is marked experimental, needs -strict -2, and produces noticeably worse quality; if a command fails with a message about the encoder being experimental, that is usually the reason. The -ar 48000 is not strictly required, since libopus accepts 8, 12, 16, 24 and 48 kHz and FFmpeg resamples anything else, but Opus always decodes to 48 kHz in Ogg, so feeding 44.1 kHz material means one resampling step either way. Doing it explicitly at encode time with a good resampler is better than leaving it to an unknown player.

Parameter Tuning Guide

  • -application voip | audio | lowdelay: voip biases the encoder toward speech intelligibility, audio (the default) toward fidelity for music and mixed content, and lowdelay disables the SILK modes so only CELT with minimal delay is used.
  • -frame_duration: 20 ms by default. 10 ms or less lowers latency but adds per-packet overhead; 40–60 ms is slightly more efficient for stored content and very low bitrates.
  • -vbr on | constrained | off: on gives the best quality per bit. constrained keeps packet sizes near the target, which some RTP systems prefer. Fully constant bitrate has a privacy argument for encrypted voice, since VBR packet sizes can leak information about what is being said.
  • -compression_level 0–10: encoder complexity (default 10). Lower values save CPU on embedded devices at a small quality cost.
  • -packet_loss and -fec 1: tell the encoder the expected loss percentage and enable in-band FEC, which only works in the SILK and hybrid modes.

Quality vs File Size Tradeoff

Mono speech is typically fine at 16–32 kbit/s, and 24 kbit/s is already clearly better than a phone call. Stereo music shows artifacts, mostly a narrowed or smeared stereo image and rough high frequencies, below roughly 96 kbit/s; around 128 kbit/s it is transparent for most listeners on most material. For an archive, start at 128 kbit/s stereo and adjust with ABX listening tests on your own content, because difficult material (harpsichord, applause, heavily compressed electronic music) exposes artifacts much earlier than typical pop.

When I first moved a spoken-word archive to Opus, the biggest mistake was not the bitrate but encoding stereo: two nearly identical microphone channels coded as stereo spend bits on a “stereo image” that does not exist. Downmixing with -ac 1 let the same bitrate go to the actual voice.


Performance Comparison

Compression Ratio vs Other Codecs

  • Low-bitrate speech: Opus beats older voice codecs such as G.711, G.722, Speex and AMR-WB at comparable bitrates, and hybrid mode gives full-band sound at bitrates where those codecs are still narrowband or wideband.
  • Music distribution: at 128 kbit/s and above, Opus, AAC-LC and Vorbis are all close to transparent, and differences depend more on encoder implementation and material than on the format. Opus’s clearest advantage is at low and medium bitrates and at low delay; at high bitrates the deciding factor is usually player support, not quality.

Encoding/Decoding Speed

libopus is written in portable C with SIMD optimizations for x86 and ARM, and decoding costs very little CPU even on phones. Encoding at complexity 10 is heavier but still real-time many times over on a single desktop core. On small embedded chips, lowering complexity is the usual lever.

Subjective Quality Assessment (MOS)

Speech quality is commonly measured with ITU-T objective models such as PESQ or POLQA, and music with MUSHRA-style listening tests. Objective scores are useful for regression testing an encoder configuration, but they do not model network behavior. For a real service, test end to end with emulated jitter and packet loss, because a configuration that scores well on a clean line can sound much worse than a lower-bitrate one with FEC enabled once loss appears.


Real-World Use Cases

Streaming Services

On-demand music services still ship AAC widely because hardware decoders and older devices support it everywhere. Opus dominates where latency matters: Discord, game voice chat, conferencing, and WebRTC-based streaming. Some large platforms also serve Opus in WebM or MP4 to browsers that support it, falling back to AAC for others.

Mobile Apps

WebRTC stacks on Android and iOS use Opus as the default audio codec. For “record and upload” features without a real-time requirement, AAC is often the pragmatic choice because the operating system provides hardware-accelerated encoders and the files play in the native player without extra work.

VoIP and WebRTC

In WebRTC, Opus is effectively required. The SDP always advertises it as opus/48000/2 regardless of the actual bandwidth or channel count, which confuses people reading SDP for the first time; the real parameters are in the a=fmtp line, for example minptime=10;useinbandfec=1. The packet time (ptime, usually 20 ms) trades latency against header overhead, and whether in-band FEC, NACK or RED is enabled decides how well speech survives a lossy mobile network.

A common real-world issue is that FEC appears enabled in SDP but does nothing, because in-band FEC is only produced when the encoder has been told to expect loss and is running in a SILK or hybrid mode. At high bitrates the encoder uses CELT-only mode and the redundancy never appears.

Browser Support

All current major browsers support Opus in WebRTC. For file playback, support differs by container: Ogg Opus, Opus in WebM and Opus in MP4 are not supported identically, and Apple platforms historically lagged. Check MDN’s codec documentation and test on the actual target devices rather than assuming a .opus file will play everywhere.


Optimization Tips

Reduce File Size While Maintaining Quality

  • For speech, start at mono 24–32 kbit/s with -application voip and go lower only after listening.
  • For music, the stereo image is usually the first thing to break when bitrate drops. ABX test carefully below 128 kbit/s.
  • Trim silence or use DTX (discontinuous transmission) in real-time use; in files, VBR already spends very few bits on silence.

Improve Encoding Speed

  • For batch jobs, running several files in parallel scales far better than tweaking encoder settings, since a single libopus encode is single-threaded.
parallel ffmpeg -y -i {} -c:a libopus -b:a 128k -ac 2 {.}.opus ::: *.wav

Batch Processing Automation

Running opusinfo (from opus-tools) on the output in CI catches truncated or malformed files before they are published. It also reports the pre-skip and the original input sample rate stored in the header.

# Validate Opus file
opusinfo output.opus
# Check for errors
if [ $? -ne 0 ]; then
    echo "Invalid Opus file"
    exit 1
fi

Common Issues and Solutions

Compatibility Issues

  • Older players and some car or smart-TV systems do not play Ogg or Opus at all. If your audience is broad, keep an AAC or MP3 rendition alongside.
  • Container confusion: the same Opus packets can be wrapped in Ogg (.opus), WebM/Matroska, MP4, or RTP. A player that supports “Opus” may support only some of these containers.
# Ogg Opus (audio only)
ffmpeg -i input.wav -c:a libopus -b:a 96k output.opus
# WebM with Opus (video + audio)
ffmpeg -i input.mp4 -c:v libvpx-vp9 -c:a libopus -b:a 128k output.webm

Timing and Duration Surprises

Opus streams begin with a pre-skip, 312 samples (6.5 ms) at 48 kHz with libopus defaults, which the decoder must discard. Tools that ignore it produce a few milliseconds of offset, which matters for gapless playback and for syncing audio with video after editing. Similarly, when you inspect an Ogg Opus file with a tool that reports “44100 Hz”, that is only the original input rate stored as metadata; the decoded output is 48 kHz.

Quality Degradation

  • Expecting music at low bitrate to sound good easily fails. Set a minimum bitrate per content type rather than one global value.
  • Game and call voice: aggressive noise suppression and automatic gain control damage the signal before the codec sees it. When users complain about “robotic” voice, check the preprocessing chain and packet loss statistics before blaming Opus.

Licensing Considerations

The reference implementation is BSD-licensed, and the main patent holders granted royalty-free licenses, which is why Opus is widely described as royalty-free. For a commercial product you still need to include the BSD license text, and, as with any codec, royalty-free grants from the contributors do not rule out claims by third parties; if your organization requires certainty, have legal review it rather than relying on blog posts, including this one.


Conclusion

Key Summary

  • Opus combines a speech codec and a music codec in one bitstream and switches between them per packet, which is why it handles mixed real-time audio so well.
  • It is the de facto audio codec for WebRTC, gaming and collaboration tools, mostly because of low delay, loss resilience and licensing.
  • It works well for files too, but choose bitrate and channel count per content type, and verify player support for the container you ship.
  • Real-time voice, video, gaming: Opus first; tune packet time, FEC and the jitter buffer together.
  • Music streaming with wide device compatibility: keep the AAC pipeline and add Opus where clients support it.
  • Podcast and voice archives: mono Opus at low bitrate saves a lot of bandwidth, if your listeners’ apps play it.

Quick Decision Guide

Need low latency? → Opus
Need maximum compatibility? → MP3 or AAC
Voice-only? → Opus (24~64 kbps)
Music with legacy support? → AAC (128+ kbps)
WebRTC/browser? → Opus (mandatory)

References

  • RFC 6716: Definition of the Opus Audio Codec
  • RFC 7845: Ogg Encapsulation for the Opus Audio Codec
  • Opus Official: https://opus-codec.org/
  • FFmpeg libopus: ffmpeg -h encoder=libopus


Frequently Asked Questions (FAQ)

Q. Which -application setting should I use with libopus in FFmpeg?

A. -application voip tunes the encoder for speech intelligibility and suits calls and voice recordings. audio, the default, targets general fidelity and is the right choice for music and mixed content. lowdelay disables the speech-oriented modes to minimize algorithmic latency, which is only worth it when latency matters more than efficiency at low bitrates.