How Video Chat Works: Encryption, Throughput, WebSockets, and WebRTC

A technical look at what happens between the moment you press “Start Call” and the moment another browser receives your video and audio.

Modern video-chat applications look deceptively simple. A browser captures a camera and microphone, compresses the resulting media, sends it across the Internet, and reconstructs it on another device.

Underneath that simple interface is a collection of protocols and systems: WebRTC for real-time media, DTLS and SRTP for encryption, ICE for connectivity establishment, STUN and TURN for NAT traversal, codecs such as VP8, VP9, H.264, and AV1, and often WebSockets for signaling.

Important distinction: WebSockets and WebRTC solve different problems. A WebSocket is generally used for application-level signaling and messaging, whereas WebRTC is designed for low-latency audio and video transport.

The Basic Architecture

A typical browser-to-browser call looks approximately like this:

┌──────────────┐ ┌──────────────┐ │ Browser A │ │ Browser B │ │ │ │ │ │ Camera/Mic │ │ Camera/Mic │ │ │ │ │ ▲ │ │ ▼ │ │ │ │ │ Encoder │ │ Decoder │ │ │ │ │ ▲ │ │ ▼ │ │ │ │ │ WebRTC │◄──── SRTP media ──────►│ WebRTC │ └──────┬───────┘ └──────▲───────┘ │ │ │ Signaling │ └────────── WebSocket ───────────────────┘ │ ▼ ┌─────────────┐ │ Signaling │ │ Server │ └─────────────┘ STUN/TURN servers assist with connectivity

The signaling server normally does not carry the actual video. Instead, it helps the two endpoints exchange information needed to establish a WebRTC connection.

Why WebSockets Are Useful

WebSockets provide a persistent, bidirectional connection between a browser and a server. Once the connection has been established, either side can send messages without repeatedly creating HTTP requests.

A signaling server can use this channel to exchange WebRTC connection information.

Client A                         Signaling Server
   │                                      │
   │──── "create-call" ─────────────────►│
   │                                      │
   │◄─── session information ─────────────│
   │                                      │
   │──── SDP offer ──────────────────────►│
   │                                      │
   │                                      │
   │                         forwards to B│
   │                                      │
   │◄──── SDP answer ─────────────────────│
   │                                      │
   │──── ICE candidate ──────────────────►│
   │                                      │
   │                         forwards to B│

Once the WebRTC connection has been established, the signaling server may no longer be involved in the media path.

This distinction can be important when designing a scalable architecture. A signaling server might handle millions of small messages while the media infrastructure handles a much larger volume of network traffic.

WebRTC and the Media Path

WebRTC is designed specifically for interactive real-time communication. It supports audio, video, data channels, congestion control, NAT traversal, encryption, and connection management.

A simplified media pipeline looks like this:

Camera │ ▼ Raw frames │ ▼ Video encoder │ ▼ Compressed frames │ ▼ RTP packets │ ▼ SRTP encryption │ ▼ UDP / network │ ▼ SRTP decryption │ ▼ RTP packets │ ▼ Video decoder │ ▼ Display

Raw video is far too large to transmit efficiently, so compression is essential.

How Much Bandwidth Does Video Require?

Uncompressed video can require enormous amounts of bandwidth. Consider 1920 × 1080 video at 30 frames per second with 24 bits per pixel:

1920 × 1080 × 30 × 24 ≈ 1.49 Gbit/s

This is before accounting for additional protocol overhead. Clearly, sending raw frames over an ordinary Internet connection is impractical.

A video codec exploits spatial and temporal redundancy to reduce this dramatically. Depending on the codec, resolution, frame rate, scene complexity, and desired quality, a 1080p stream might instead operate in the range of a few megabits per second.

Video Example Bitrate Approx. Data / Hour
360p 0.5 Mbps ~225 MB
720p 1.5 Mbps ~675 MB
1080p 3 Mbps ~1.35 GB
1080p high quality 6 Mbps ~2.7 GB

These are illustrative calculations rather than fixed codec requirements. Real WebRTC sessions dynamically adjust their bitrate.

Throughput Is Not the Same as Latency

Video conferencing requires both sufficient throughput and low latency. A connection can have enormous bandwidth while still producing poor interactive video if packets take too long to arrive.

Several properties matter:

Interactive media generally benefits more from low latency than from extremely high peak bandwidth.

Why UDP Is Commonly Used

Real-time media commonly uses UDP because it avoids some of the retransmission behavior associated with TCP.

Suppose packet 100 is lost while packets 101 through 105 arrive. With a latency-sensitive video stream, waiting for packet 100 to be retransmitted may be worse than simply allowing the decoder to conceal the missing information.

This is one reason WebRTC uses RTP-based media transport with mechanisms designed for real-time conditions.

Tradeoff: UDP does not magically provide reliable or congestion-free communication. WebRTC implements additional mechanisms for congestion control, packet loss handling, retransmission where appropriate, and adaptation.

Encryption

WebRTC media is encrypted. The media path uses SRTP, while DTLS is used to establish cryptographic material for the secure transport.

Conceptually:

Application │ ▼ Compressed media │ ▼ RTP │ ▼ SRTP encryption │ ▼ Network packets

The encryption process protects the media while it travels across the network. WebRTC also authenticates the cryptographic negotiation so that endpoints can establish a secure association.

Encryption vs. End-to-End Encryption

It is important to distinguish transport encryption from application-level end-to-end encryption.

In a direct peer-to-peer connection, encrypted media can travel directly between participants. However, conferencing architectures frequently use a media server or Selective Forwarding Unit (SFU).

An SFU receives media packets and forwards them to other participants. Depending on the architecture and cryptographic design, this can have different implications for who can access decrypted media.

NAT Traversal: STUN and TURN

Most devices are not directly reachable from the public Internet. Home routers, corporate firewalls, and carrier networks commonly use NAT.

WebRTC uses ICE (Interactive Connectivity Establishment) to discover possible network paths.

STUN

A STUN server can help a device determine its public-facing network address and port.

TURN

Sometimes a direct peer-to-peer connection is impossible. A TURN server can then relay the media between the participants.

Direct connection: Browser A ─────────────────────── Browser B Relayed connection: Browser A ─────► TURN Server ─────► Browser B │ │ encrypted media

TURN has an important infrastructure consequence: the server now carries the media traffic. A service supporting many calls can therefore require substantial network bandwidth even if the application has relatively little signaling traffic.

Peer-to-Peer vs. SFU Architecture

A two-person call can potentially use a direct peer-to-peer connection. Group calls become more complicated.

Mesh

In a naive mesh architecture, every participant sends a separate stream to every other participant.

Upload streams per participant = N − 1

Consequently, a ten-person conference could require each participant to upload nine separate streams.

SFU

An SFU allows each participant to upload a stream once. The server then forwards appropriate streams to the other participants.

┌──────────────┐ │ SFU │ └──────┬───────┘ │ ┌────────────┼────────────┐ ▼ ▼ ▼ Browser A Browser B Browser C

Modern conferencing systems frequently use this architecture because it makes upload requirements much more manageable.

Adaptive Bitrate

A video-chat system cannot assume that the network will remain constant. A participant might move from Wi-Fi to cellular data, encounter congestion, or experience packet loss.

WebRTC can react by changing the media bitrate and other encoding parameters.

For example, a sender might transition approximately like this:

Good network │ ▼ 1080p @ 30 FPS │ │ congestion detected ▼ 720p @ 30 FPS │ │ bandwidth decreases ▼ 480p @ 24 FPS │ │ severe congestion ▼ 360p @ 15 FPS

The exact behavior depends on the browser, codec, congestion controller, network conditions, and application configuration.

Video Codecs

The codec determines how raw camera frames are transformed into a compressed bitstream.

Codec General Characteristics
VP8 Widely supported and relatively straightforward to encode.
VP9 More compression efficiency, with greater computational complexity.
H.264 Broad hardware support and extensive ecosystem compatibility.
AV1 High compression efficiency, with potentially significant encoding complexity.

Hardware acceleration is particularly important for video chat. Encoding and decoding high-resolution video entirely on the CPU can consume substantial computational resources.

Audio Is Different

Audio usually requires dramatically less bandwidth than video, but conversational audio is extremely sensitive to latency.

Codecs such as Opus are commonly used because they can operate across a wide range of bitrates while supporting interactive communication.

A video call can therefore continue to provide intelligible speech even when video quality has been substantially reduced.

A Minimal Browser Example

The browser API for obtaining camera and microphone access is straightforward:

const stream = await navigator.mediaDevices.getUserMedia({
    video: true,
    audio: true
});

videoElement.srcObject = stream;

A WebRTC peer connection can then be created:

const pc = new RTCPeerConnection({
    iceServers: [
        { urls: "stun:stun.example.com" }
    ]
});

for (const track of stream.getTracks()) {
    pc.addTrack(track, stream);
}

The application still needs signaling code to exchange offers, answers, and ICE candidates with the other endpoint.

Scaling a Video-Chat Service

Scaling video chat is fundamentally different from scaling a normal web application.

A conventional web server might process thousands of small HTTP requests. A video infrastructure system can instead maintain large numbers of persistent, high-bandwidth media flows.

Suppose an SFU forwards 2 Mbps to and from each participant. With 10,000 simultaneously connected participants, even a simplified calculation can produce tens of gigabits per second of aggregate traffic.

Consequently, production systems must consider:

The Key Takeaway

A modern browser video call is not simply “video sent through a socket.” It is a layered real-time communication system.

┌─────────────────────────────────────┐ │ Video application │ ├─────────────────────────────────────┤ │ WebRTC / media management │ ├─────────────────────────────────────┤ │ RTP / SRTP / DTLS │ ├─────────────────────────────────────┤ │ ICE / STUN / TURN │ ├─────────────────────────────────────┤ │ UDP / IP / network │ └─────────────────────────────────────┘ WebSocket │ └── Signaling, presence, chat, session negotiation, etc.

WebSockets are well suited to coordinating a call, while WebRTC handles the difficult real-time media problem. Encryption protects media in transit, codecs reduce the enormous bandwidth requirements of raw video, and adaptive transport mechanisms allow the system to respond to changing network conditions.

Once these layers are understood independently, the architecture of seemingly complex applications such as video conferencing becomes considerably easier to reason about.