Video Call Quality Test: Latency, Jitter, and Loss

WebRTC peer-to-peer latency, jitter, and packet loss. All measurement runs in the browser. No backend, no account.

Video Call Quality Test: Latency, Jitter, and Loss

Three metrics determine video call quality. Round-trip time sets the delay between what one participant says and when the other hears it; 150 ms is the threshold above which callers feel the conversation is lagging1. Jitter determines whether that delay is consistent or erratic; 20 ms of jitter forces the application's jitter buffer to add a matching fixed delay to smooth out arrival variation2. Packet loss drops frames and audio samples; above 3%, most video codecs cannot conceal the gaps and visible artifacts appear3.

Measuring all three over a direct peer-to-peer WebRTC path gives you the real quality numbers your call application experiences, not an idealized server-to-client measurement. Corporate firewalls, symmetric NAT, and ISP congestion affect the P2P path in ways that a server ping does not reveal. This tool establishes a WebRTC data channel between two browser instances and runs continuous probe traffic so all three metrics stabilize over a 30-second window.

What to look for

  • RTT under 150 ms, jitter under 10 ms, loss 0%
  • RTT under 80 ms, jitter under 5 ms, loss under 0.5%

Opens the P2P Network Tester with this section's reference values shown at the top of the tool.

Open in the tool →

How RTT, jitter, and loss interact in video calls

High RTT without jitter produces a noticeable conversation delay but intelligible audio. Both participants adapt to the fixed delay within a few exchanges. High jitter with moderate RTT is more disruptive; the jitter buffer must grow to absorb timing variation, adding its own fixed delay on top of the network RTT. Packet loss compounds both effects: the codec concealment algorithm works well up to 3% loss but degrades quickly above that3, and error correction codecs add redundant data that further increases bandwidth consumption. Consequently, 80 ms RTT with 5 ms jitter and 0.5% loss produces a far better call experience than 50 ms RTT with 30 ms jitter and 2% loss.

Why jitter and loss dominate perceived call quality

Jitter and packet loss together have a larger impact on perceived call quality than raw RTT at moderate latency levels. A stable 150 ms connection with 2 ms jitter and 0% loss produces better call quality than an unstable 50 ms connection with 30 ms jitter and 2% loss, even though the second connection has lower latency. The reason is that jitter and loss cause audible clicks, video freezes, and codec degradation, while a fixed 150 ms delay is something callers adapt to within a few exchanges.

Quality benchmarks for video call applications

WebRTC-based video calls typically perform well within these bounds: RTT below 150 ms, jitter below 20 ms, and packet loss below 1%. Zoom, Google Meet, and Microsoft Teams publish similar thresholds in their network requirement documentation4. Building on these thresholds, a connection with RTT of 100 ms but 0% loss and 3 ms jitter will perform better than a connection with 60 ms RTT, 25 ms jitter, and 2% loss, even though the second connection has lower latency. Jitter and loss together matter more than raw RTT for perceived call quality at these moderate latency levels. CapyToolkit reports all three metrics simultaneously with the same thresholds these platforms publish, so you can compare your readings directly against the published requirements for whichever video conferencing platform you rely on.

Using the simulation sliders to verify your application

The loss and latency simulation sliders let you introduce controlled degradation to your P2P path. Slide the loss rate to 3% and observe whether your video call application handles the drops gracefully or begins producing visual artifacts. Add 100 ms artificial latency and verify the conversation echo behavior. Furthermore, combining 5% simulated loss with 50 ms artificial latency replicates a congested mobile network, and testing this combination reveals whether your application adapts its codec or drops to audio-only mode under those conditions. Running a five-minute test under combined 3% loss and 80 ms latency simulates a congested residential ISP path during peak hours, which is the most common real-world video call degradation scenario that home users experience when other household members are streaming or downloading in the background.

How bandwidth saturation affects RTT, jitter, and loss simultaneously

Bandwidth saturation elevates all three quality metrics at once because the mechanism is queue overflow. When upload capacity is fully consumed by a competing transfer, the router delays or drops packets from all outbound streams uniformly5. RTT rises as probe packets wait in the full queue; jitter increases because queue depth fluctuates as packets enter and leave; packet loss spikes when the queue overflows and the router begins dropping rather than delaying. Recognizing this three-metric simultaneous rise identifies upload saturation as the cause rather than an ISP path problem.

Uploading a large file saturates most residential upload links and causes all three quality metrics to spike together. The fix is either closing the competing upload or enabling QoS on the router to prioritize real-time UDP. Testing with and without a competing upload confirms the diagnosis: all three metrics return to baseline when the background transfer stops.

Bandwidth requirements for common video call configurations

Google Meet requires 3.2 Mbps upload for HD 720p and 4 Mbps for 1080p group video6. Microsoft Teams requires 1.5 Mbps per participant for group HD video7. Zoom requires 3.8 Mbps for group HD4. These are per-session upload figures; on a 20 Mbps upload connection you have headroom for one high-quality video call with several Mbps remaining for background traffic. Enabling QoS and reserving 10 Mbps for video call UDP traffic while limiting background transfers prevents saturation-driven quality degradation even when other household members are active.

Platform network requirements and compliance testing

Zoom, Google Meet, and Microsoft Teams all publish minimum and recommended network requirements in their official support documentation. Zoom specifies 600 kbps for 360p, 1.2 Mbps for 720p HD, and 3.8 Mbps for group HD. Google Meet specifies 3.2 Mbps for HD 720p and 4 Mbps for 1080p in group calls. Microsoft Teams specifies 1.5 Mbps for full-screen HD video. All three platforms share similar latency and jitter thresholds: RTT below 150 ms, jitter below 20 ms, and packet loss below 1% for consistently high-quality calls.

Verifying compliance before a deployment means measuring all three metrics using this tool during a representative network load window. Running the test during expected peak-use hours, with typical background applications active, reveals whether the connection meets thresholds under realistic conditions rather than during off-peak idle periods. CapyToolkit does not enforce any platform-specific thresholds; it reports the raw RTT, jitter, and loss numbers so you can compare them against whichever platform's requirements you need to meet.

Documenting network conditions for IT support escalation

When reporting call quality problems to IT support or an ISP, screenshots of this tool's metric display with timestamps provide concrete diagnostic evidence. Capture the sparkline, average RTT, jitter, and loss percentage during the problem window. Note the ICE candidate type; "relay" suggests the network is forcing relay routing that adds latency beyond the platform's design assumptions. This documentation distinguishes network-layer problems from application or hardware problems before the escalation begins.

Attaching the timestamps turns a subjective complaint into a measurable record that support staff can correlate with their own logs on the same connection. A sparkline showing relay routing during the problem window is far more persuasive than describing the call as choppy after the fact. Keeping a short history of these captures across several sessions also reveals whether the issue is specific to one meeting platform or present on every call over the same network path.

When to use this

Run this test before an important video call to verify your network path, when callers report audio or video quality problems you cannot reproduce in a speed test, or when you are matching your ISP link to Zoom's bandwidth minimums before committing to a new video conferencing platform.

Examples

Call quality fine for 10 minutes then degrades

Jitter spikes from 4 ms to 35 ms at a regular interval. This pattern suggests a router applying traffic shaping after a bandwidth budget is exceeded, or a background application starting a scheduled download. Identify the competing process and restrict its bandwidth.

Remote participant reports echo; local participant hears nothing unusual

Echo is typically an audio feedback issue at one endpoint, not a network quality problem. Network quality metrics appear normal. The issue is the remote participant's microphone picking up speaker output. Ask them to use headphones or reduce speaker volume.

Sources
  1. 1.

    ITU-T, "One-way transmission time," Recommendation G.114, International Telecommunication Union, May 2003. https://www.itu.int/rec/T-REC-G.114/en

  2. 2.

    Twilio, "Improve Call Experience with New Twilio Conference Jitter Buffer Controls," twilio.com, June 2020. https://www.twilio.com/en-us/blog/products/launches/improve-call-experience-new-twilio-conference-jitter-buffer-controls

  3. 3.

    "Twilio Voice Insights — SDK Call Quality Events," Twilio, accessed June 2026. https://static0.twilio.com/docs/voice/voice-insights/api/call/details-sdk-call-quality-events

  4. 4.

    "Zoom system requirements (Windows, macOS, Linux)," Zoom, accessed June 2026. https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0060748

  5. 5.

    J. Gettys, "Bufferbloat: What's Wrong with the Internet?," ACM Queue, vol. 9, no. 11, Association for Computing Machinery, November 2011. https://dl.acm.org/doi/10.1145/2063166.2071893

  6. 6.

    "Google Meet hardware requirements," Google Workspace, accessed June 2026. https://support.google.com/a/answer/4541234

  7. 7.

    "Monitor call and meeting quality in Microsoft Teams," Microsoft, accessed June 2026. https://support.microsoft.com/en-us/teams/meetings/monitor-call-and-meeting-quality-in-microsoft-teams

VoIP Quality Test: Latency, Jitter, and Packet Loss

VoIP quality degrades before users notice why. A 2% packet loss rate translates to roughly one dropped word per sentence in a standard 50-word-per-minute conversation, audible but not always attributed to the network. At 30 ms jitter, the receiver's playout buffer must add 30 ms of fixed latency to smooth out arrival variation, often pushing end-to-end delay above the 150 ms threshold for natural conversation1. Understanding these three metrics and how they combine gives you the diagnostic data to identify the actual cause of a poor call before blaming the codec or application.

This tool measures RTT, jitter, and packet loss over a direct peer-to-peer WebRTC path using probe packets between two browser instances. The path mirrors the one your VoIP application uses when it connects directly between participants, revealing the real network conditions rather than a cleaner server-to-client baseline.

What to look for

  • below 1%
  • above 3%
  • 150 ms (about 300 ms RTT)

Opens the P2P Network Tester with this section's reference values shown at the top of the tool.

Open in the tool →

ITU-T G.114 thresholds for VoIP quality

ITU-T G.114 defines 150 ms one-way delay as the threshold for interactive voice; above this, callers notice the gap between speaking and receiving a response. Round-trip time in this tool is approximately twice the one-way delay, so 300 ms RTT represents the ITU-T boundary1. G.114 also notes that above 400 ms RTT, the quality of interaction degrades significantly regardless of codec or application. Furthermore, the ITU-T E-model (R-value) combines delay, jitter, and packet loss into a single quality score: R above 70 is acceptable, above 80 is good, and 90 or higher is toll-quality audio2. CapyToolkit reports the same RTT, jitter, and loss values the E-model uses, so you can compare your readings against these thresholds without needing a separate E-model calculator.

How the E-model combines RTT, jitter, and loss into one score

The E-model treats delay, jitter, and packet loss as independent degradations that subtract from a maximum achievable quality score. At 300 ms RTT with 20 ms jitter and 1% loss, the R-value lands near 70, which is the minimum acceptable threshold for business use. Reducing any single metric raises the score: cutting jitter from 20 ms to 5 ms at the same RTT adds roughly 10 points to the R-value, making improvement efforts targeting the most degraded metric the most effective use of troubleshooting time.

How VoIP codecs handle jitter and loss

VoIP codecs use adaptive jitter buffers to absorb timing variation. The buffer holds incoming audio packets temporarily and replays them at a fixed rate, smoothing out arrival jitter at the cost of added fixed delay. A jitter buffer configured for 40 ms of jitter adds 40 ms of delay to every call, regardless of what the actual jitter is during that interval. Building on this, packet concealment algorithms (built into all modern codecs including Opus, G.711, and G.729) generate synthetic audio for lost packets up to about 3%3. Above that threshold, concealment sounds unnatural and callers hear audible dropouts. The codec's built-in forward error correction (FEC) adds redundant data alongside primary audio packets, allowing the receiver to reconstruct some lost packets without perceptible quality loss at moderate loss rates up to 2 to 3 percent.

Diagnosing VoIP problems using P2P metrics

A VoIP call that sounds distant or delayed maps to high RTT; check the average RTT in this tool. Clipping or choppy audio maps to high jitter; check both average jitter and the max jitter spike visible in the sparkline. Audio that drops out in short bursts maps to packet loss; check the real loss percentage. Conversely, audio with distortion or compression artifacts is a codec or bandwidth issue, not a network metric problem; this tool does not measure audio bandwidth. Identifying which metric is out of range narrows the diagnosis to a specific layer before attempting any fix. CapyToolkit reports the three network inputs to call quality (RTT, jitter, and loss) as separate values, so you can determine which one is the dominant problem before changing any router or ISP settings.

Codec selection and DSCP marking for VoIP prioritization

VoIP codec choice sets the quality ceiling for any given network path. Opus, the dominant codec in WebRTC and modern SIP deployments, adapts bitrate in real time from 6 kbps to 510 kbps based on detected network conditions3. G.711 (PCM audio), the standard codec for enterprise SIP trunks and PSTN connections, uses a fixed 64 kbps payload with no adaptive bitrate or forward error correction4. Because G.711 transmits no redundancy, it is more sensitive to packet loss than Opus: 2% loss produces audible gaps in G.711 while Opus's built-in FEC recovers most lost frames at the same loss rate. When evaluating this tool's loss readings, the acceptable threshold differs by codec.

DSCP tagging classifies outbound RTP audio packets for router-level QoS prioritization. The standard DSCP value for VoIP media is EF (Expedited Forwarding, DSCP 46)5; routers recognizing this marking place tagged packets in the highest-priority queue ahead of HTTP downloads and best-effort traffic. Most enterprise routers and gaming routers honor DSCP markings on the WAN interface. Applying DSCP 46 to VoIP audio prevents competing file transfers from delaying audio packets in the upload queue, reducing jitter without requiring per-port traffic shaping rules.

Configuring DSCP marking on common platforms

On Windows, the Group Policy DSCP marking policy (under Computer Configuration > Windows Settings > Policy-based QoS) applies DSCP 46 to outbound traffic from a specific application. Cisco IP Communicator, Zoom Phone, and Teams VoIP all support DSCP configuration through enterprise management interfaces. For home VoIP adapters, DSCP marking is typically configured per-port in the ATA (Analog Telephone Adapter) settings panel, often labeled "QoS" or "DSCP marking" under the device's web interface.

One-way audio failures and asymmetric path diagnostics

One-way audio means one participant hears the other clearly but their own voice does not reach the far end. This failure almost never appears in symmetric RTT measurements because probe packets complete the round trip successfully; the asymmetry is in the application-layer audio path, not the probe path. The most common cause is a SIP ALG (Application Layer Gateway) on the router rewriting SDP port numbers in the SIP INVITE message, creating a mismatch between the negotiated RTP port and the port the application actually uses6.

Disabling SIP ALG on the router resolves most one-way audio issues on consumer hardware. Find this setting under Firewall or WAN settings labeled as SIP ALG, SIP Passthrough, or SIP Helper depending on the firmware. After disabling SIP ALG, disconnect and reconnect the VoIP client to force a fresh SDP negotiation. If one-way audio persists after disabling SIP ALG, a firewall rule is blocking inbound UDP on the RTP port range and requires a port forwarding or allow rule for the remote endpoint's IP address.

Identifying asymmetric path blocks with this tool

Swapping sender and receiver roles mid-test starts with your device as the probe sender; note the loss counter, then reverse roles (the other browser becomes the sender) and note the second loss counter. Matching loss rates confirm a symmetric path; a loss rate appearing only in one direction confirms an asymmetric block. For VoIP troubleshooting, if your instance shows 0% loss as sender but 100% loss as receiver, inbound UDP on the return port is blocked at the router level, isolating the issue before any ISP escalation begins.

Running the two roles back to back on the same session isolates the direction of the block, which a single one-directional test would completely miss. An inbound-only drop almost always means a firewall or NAT rule is rejecting return UDP, since outbound packets clearly reached the peer. Capturing both counters before contacting a provider prevents the common misdiagnosis of blaming codec quality when the real fault is an asymmetric filter on the return path that no amount of bandwidth will fix.

When to use this

Run this test when VoIP call quality is consistently poor, when callers report choppy audio that a speed test cannot explain, or when deploying a VoIP system and verifying that the network path meets ITU-T G.114 requirements.

Examples

One-way audio on incoming calls, outgoing audio works fine

One-way audio is almost always a NAT or firewall issue preventing return audio packets from reaching one endpoint, not a network quality problem. The P2P metrics appear normal because the probe packets succeed. Check the NAT type and ensure the VoIP port is open bidirectionally.

1.5% real packet loss causing choppy audio

The loss percentage appears in the tool's counters. The audio codec is concealing the loss but the gaps are audible. Identify the loss source: run the test during the same time window when calls are choppy and compare morning versus evening loss rates to determine whether the cause is ISP congestion.

Sources
  1. 1.

    ITU-T, "One-way Transmission Time," Recommendation ITU-T G.114, International Telecommunication Union, May 2003. https://www.itu.int/rec/T-REC-G.114-200305-I/en

  2. 2.

    ITU-T, "The E-model: a Computational Model for Use in Transmission Planning," Recommendation ITU-T G.107, International Telecommunication Union, June 2015. https://www.itu.int/rec/T-REC-G.107-201506-I/en

  3. 3.

    J. Valin, K. Vos, and T. Terriberry, "Definition of the Opus Audio Codec," RFC 6716, IETF, September 2012. https://www.rfc-editor.org/rfc/rfc6716.txt

  4. 4.

    "G.711," Wikipedia, accessed June 2026. https://en.wikipedia.org/wiki/G.711

  5. 5.

    S. Blake, D. Black, M. Carlson, E. Davies, Z. Wang, and W. Weiss, "An Architecture for Differentiated Services," RFC 3246, IETF, December 2002. https://www.rfc-editor.org/rfc/rfc3246.txt

  6. 6.

    "Application-level gateway," Wikipedia, accessed June 2026. https://en.wikipedia.org/wiki/Application-level_gateway

FAQ

CapyToolkit reports real packet loss from timed-out probes rather than inferred loss, giving you the actual loss figure your VoIP codec would experience on the same path. Below 1% is ideal. Most codecs with packet concealment maintain acceptable quality up to 3%. Above 3%, audible dropout occurs in nearly all codec implementations. At 5%, callers will consistently report call quality problems.

This tool measures round-trip time (the time for a probe to travel to the peer and return). One-way VoIP latency is approximately half the RTT. The ITU-T G.114 threshold is 150 ms one-way, which corresponds to roughly 300 ms RTT in this tool.

This tool measures the WebRTC path, not the SIP trunk path. However, if the P2P path to your SIP gateway location shows good metrics here, the SIP trunk path will likely perform similarly. For SIP trunk qualification, use a dedicated SIP testing tool that sends actual INVITE messages and RTP audio packets.

Time-of-day quality variation typically indicates ISP congestion during peak usage hours. Run this tool during both the good period and the bad period and compare the RTT, jitter, and loss readings. If jitter and loss increase in the afternoon, contact your ISP with the timed measurements to document the congestion.

No. Calculating MOS requires analyzing actual audio samples or running an ITU-T E-model computation on call metrics this tool does not collect. This tool measures the three network inputs to MOS: RTT, jitter, and packet loss. Use a dedicated VoIP quality analyzer to compute MOS from these inputs.

FAQ

CapyToolkit displays all three metrics simultaneously so you can identify which one is the dominant problem on your connection. Jitter and packet loss together have the most impact on perceived quality. RTT matters for conversation feel, but jitter and loss directly cause audio clicks, video freezes, and codec degradation. A stable 150 ms connection with 2 ms jitter and 0% loss produces better call quality than an unstable 50 ms connection with 30 ms jitter.

Yes. Run the test 10 to 15 minutes before the interview from the same network and device you plan to use for the call. If RTT is below 150 ms, jitter is below 10 ms, and loss is 0%, your network is ready. If any metric is marginal, switch to a wired connection or close bandwidth-consuming applications.

Frame drops in a video call can originate from CPU encoding limitations, insufficient bandwidth for the selected video resolution, or the codec hitting its complexity ceiling, none of which appear as packet loss in this tester. Check CPU usage and available bandwidth during the call. Lowering the camera resolution often eliminates frame drops caused by encoding overhead.

No. The probe packets travel in both directions simultaneously. The round-trip measurement captures the combined upstream and downstream path. Asymmetric congestion (where upload is clean but download is saturated) appears as elevated RTT and loss from the perspective of the peer whose download path is congested.

Target RTT below 80 ms, jitter below 5 ms, and packet loss below 0.5% for a consistently high-quality experience. Most wired home connections on a residential ISP achieve these numbers outside peak hours. Evening congestion on shared neighborhood infrastructure can push RTT and loss above these thresholds.

Additional resources