AudioMundo

A plain-language reference desk for sound, audio formats, and levels

AES67 and SMPTE ST 2110-30: How Do They Fit Together?

AES67 is an Audio Engineering Society standard for carrying uncompressed PCM audio over ordinary IP networks in a way that different manufacturers' equipment can interoperate. SMPTE ST 2110-30 is the audio part of the SMPTE suite for professional media over managed IP networks, and it carries PCM audio by reference to AES67 rather than defining a separate transport. In practice a device described as "ST 2110-30" is sending an AES67-compatible audio stream inside a larger ST 2110 system.

What each standard covers

  • AES67 defines the audio stream: RTP packetization, clock synchronization, and the baseline media formats a compliant device must handle. It is published through the AES standards program, alongside AES3 and the rest of the AES catalog.
  • SMPTE ST 2110 is a suite, not one document. ST 2110-20 carries uncompressed video, ST 2110-30 carries PCM audio, and ST 2110-40 carries ancillary data — each as a separate elementary stream on the same network. SMPTE's own ST 2110 overview describes the suite and its parts.

The practical consequence of separate elementary streams is that audio, video, and data are routed independently. There is no embedded audio to de-embed; an audio device subscribes to an audio stream and never touches the video.

The role of the clock

Both rely on a shared network clock rather than a word clock distributed on coax. Precision Time Protocol (IEEE 1588 profiles) distributes time across the network, and every sender timestamps its packets against it. This is the main conceptual break from cable-based digital audio: in an AES3 link the clock rides on the signal itself, so the cable carries both data and timing. On an IP network, timing arrives separately, and a receiver that is not locked to the same grandmaster clock will not align with anything else.

Sampling rate and payload

AES67 interoperability is built around 48 kHz PCM. The sampling rate is simply the number of samples taken per second per channel — the federal digitization guidelines define it that way, with the Nyquist–Shannon basis noted (FADGI glossary: sampling rate). Combined with a word length of 16 or 24 bits, that determines the payload rate directly:

  • 48,000 samples/s × 24 bits × 2 channels = 2,304,000 bits/s, or about 2.3 Mbps of audio payload for a stereo stream.
  • 48,000 × 24 × 8 = 9,216,000 bits/s for an eight-channel stream.

RTP, UDP, IP, and Ethernet headers add to that, and the overhead fraction grows as packet time shrinks: shorter packets mean fewer audio samples per header. Packet time is therefore a latency-versus-overhead trade-off, not a quality setting — the samples are uncompressed either way. The same arithmetic applied to file storage is covered in how to calculate audio bit rate and file size.

What this means on a spec sheet

  • "AES67" on its own says the device can exchange uncompressed audio streams with other AES67 devices, given a common clock and a network configured for it.
  • "ST 2110-30" says the device fits into a full SMPTE IP media plant, where its audio is one stream among video and ancillary data streams.
  • "AES3" describes a point-to-point cable interface — two channels down one physical link — and is not an IP technology at all. The differences between these labels are laid out in what AES3, AES67, and BS.1770 mean on a spec sheet.

Many products list all three, because the interfaces coexist: a converter may accept AES3 on XLR and emit an AES67 stream on an RJ45 port.

Where to look next

Profiles, constraints, and conformance details change between editions, and vendor summaries lag. The AES standards catalog and the SMPTE ST 2110 overview linked above are the authoritative places to check what the current editions actually require.

Sources