Claude Voice Mode Mic Test

Test your microphone before Claude Voice Mode sessions. Clipping causes transcript errors. Check noise floor below −50 dBFS before any voice conversation.

ZERO UPLOAD · ALL LOCAL
  1. Click "Enable Microphone" and allow access in the browser prompt — microphone access is only used locally for analysis.
  2. Select a test from Room Acoustics: Noise Floor Grade, Clipping Detector, or Frequency Response.
  3. Select a test from Playback & Latency: Echo Loopback or Clap Latency Test.
  4. Noise Floor: stay completely silent, click "Start 3-second test", and read the dBFS result and grade.
  5. Clipping Detector: speak at normal volume; watch for the red CLIPPING badge — reduce your input gain if it appears.
  6. Frequency Response: speak or play audio continuously; observe the live FFT spectrum across 20 Hz–20 kHz.
  7. Echo Loopback: click "Record & Play Back" and listen to the 3-second playback for echo or quality issues.
  8. Clap Latency: wear headphones, click "Start Listening", then clap once sharply near the microphone.

What to look for

  • above -40 dBFS noise floor

Microphone access is required to run any test. Access is only used for analysis — never recorded or transmitted.

Microphone active — select a test below
Room Acoustics
Playback & Latency

Stay completely silent, then click Start to measure your room's background noise level over 3 seconds.

— dBFS CLIPPING

Weak energy below 200 Hz = thin-sounding mic. Weak energy above 4 kHz = muffled audio.

Click to record 3 seconds and hear playback through your speakers.

Click Start Listening, then clap once sharply near your microphone. Use headphones to prevent feedback.

Includes speaker output, room travel, and mic input. Typical browser audio stack: 20–80 ms.

Claude Voice Mode Mic Test: Noise Floor, Clipping and Latency Check

Claude Voice Mode depends on a speech-to-text stage before response generation. Unlike end-to-end audio models, this pipeline has distinct stages: audio capture, speech recognition, and language model inference.1 Clipping at the audio capture stage introduces distortion that corrupts the speech recognition stage specifically: the waveform shape the recognition model depends on is destroyed by digital saturation, producing transcript errors that the language model then tries to interpret from damaged input.2

A noise floor below −50 dBFS is the threshold where speech recognition errors from background noise become negligible in Claude's voice pipeline for standard accents and speaking rates. Above that level, error rates in the transcription stage increase gradually, though they rarely become catastrophic until the floor exceeds −40 dBFS. The Noise Floor Grade, Clipping Detector, and Frequency Response display are the three tests to run before a voice session. Specifically, ensure the Clipping Detector shows no badge during your loudest normal speech, and that the Noise Floor Grade reads Good or Excellent.

Noise floor requirements

Claude's voice pipeline transcribes audio before the language model processes it. Background noise that sits between −50 dBFS and −40 dBFS (the Noisy range on the Noise Floor Grade) reduces transcription accuracy by competing with the speech signal in critical frequency bands.3 Consonants, particularly the sibilant frequencies between 3–8 kHz, are most affected because their energy is weaker relative to fundamental vowel frequencies. Furthermore, sustained background noise causes the speech recognition model to allocate computational attention to modeling noise rather than speech, which reduces confidence scores on phoneme-to-word mapping.

Sibilant sounds like "s," "f," and "sh" produce their distinguishing acoustic energy in the 3 to 8 kHz range, where amplitude is naturally 15 to 25 dB below the fundamental vowel energy concentrated below 1 kHz.4 When background noise fills this upper frequency band, the speech recognition model receives a signal where the noise floor and the sibilant energy occupy the same amplitude range, making it impossible to distinguish between a genuine "s" sound and a noise burst that happens to have similar spectral character. Running the Noise Floor Grade before each session catches this problem before it manifests as repeated transcription errors on the most common consonants in English.

Latency and interruption handling

Claude Voice Mode responds after your speech turn ends, using silence detection to identify turn boundaries. Round-trip latency measured by the Clap Latency Test affects the delay you perceive between asking a question and hearing Claude begin its response. Latency below 60ms keeps the browser-side audio stack out of the critical path; at that level, network and server processing time dominate the perceived delay. Building on this, if the Echo Loopback test reveals significant distortion in the playback, investigate whether your monitoring setup is creating acoustic feedback into the microphone during Claude's audio output.

Claude's turn-taking depends on detecting a sufficient drop in audio amplitude after you finish speaking, which signals that the model should begin generating its response. When browser-side latency is high, the silence detection window at the server receives your audio delayed by the full round-trip time, which means the model waits longer than necessary before responding and the conversation develops an unnatural rhythm where each turn begins with a perceptible pause. Keeping the Clap Latency Test result in the Good range or better minimizes this browser-side contribution and lets the server-side silence detection operate on the most current audio available.

Clipping and frequency response

Clipping produces harmonic distortion that extends across the entire spectrum above the clipped frequency.2 For Claude's speech recognition pipeline, clipping during peak speech levels, particularly during plosive consonants like p and b, corrupts the short audio frames the acoustic model processes. The Clipping Detector triggers when more than 1% of samples in a frame hit 0 dBFS. Even brief episodes at this level during normal speech indicate gain staging that is too high. The Frequency Response display helps confirm that the gain setting that eliminates clipping still delivers sufficient energy through the 200 Hz–4 kHz fundamental speech range.

The language model stage that follows transcription can often infer the correct word from surrounding context even when the transcript contains minor errors, but the speech recognition stage that produces the initial transcript has no such contextual safety net for individual phonemes. A clipped "p" sound at the start of a word produces a distorted waveform that the acoustic model maps to an entirely wrong phoneme candidate, and the language model downstream receives this wrong phoneme sequence with no way to know the original signal was corrupted by digital saturation rather than being a genuine speech sound. This is why the Clipping Detector matters more for Claude Voice Mode than for text-based interactions.

Gain staging for Claude's speech recognition pipeline

Because Claude's voice pipeline uses a speech recognition stage before the language model processes your input, gain staging affects transcript accuracy directly rather than indirectly through a processing buffer. Starting at 60% OS input gain provides enough signal for a usable Noise Floor Grade in most quiet rooms while leaving headroom below the Clipping Detector threshold for a typical voice. Run the Noise Floor Grade first; if the result is Noisy, increase gain by 5% and repeat. If the result is Good or Excellent, run the Clipping Detector while speaking at maximum normal volume before beginning any session.

Finding the optimal gain bracket in under five steps

The search for the optimal gain setting takes five adjustment steps at most in a typical room. Start at 40% OS input gain, run the Noise Floor Grade, and note the result. Increase to 50% and repeat. At each step, the grade improves until room noise becomes the dominant source and further gain increases stop improving it. That inflection point is the optimal gain setting. Confirm by running the Clipping Detector at that gain; if no badge appears during loud speech, the setting is both noise-floor-optimal and clipping-safe.

Confirming the safe window before each session

After finding the gain level where the Noise Floor Grade is Good and the Clipping Detector shows no badge during loud speech, document the OS gain slider position with a screenshot. Revisit this calibration when you change rooms, change hardware, or reconnect the microphone after a period of non-use. Gain staging that was optimal in a previous session may be too high if the OS default communications device changed and selected a different microphone at its default gain setting. The two-minute dual-test sequence prevents this common source of degraded voice sessions.

The dual-test sequence is the Noise Floor Grade followed immediately by the Clipping Detector at the same gain, completed in under two minutes. Keeping the screenshot of the OS slider beside the microphone means you can restore the exact position after any device change without re-running the full search, which is the practical value of documenting the calibration rather than relying on memory of a percentage.

Understanding how soft speech affects clipping risk

Clipping risk in Claude Voice Mode is determined by your loudest typical speech, not your average or softest speech. The Clipping Detector threshold of 1% samples at 0 dBFS means it fires when any speech peak exceeds the digital ceiling, regardless of how quiet the rest of the recording is. If your gain staging is calibrated to your whispered or quiet voice, the gain may be set too high to handle emphatic speech, technical explanations delivered with more force, or the natural loudness increase that occurs when speaking about something important.

Calibrating with loud consonants as the test signal

During the Clipping Detector calibration run, include deliberate emphatic speech rather than only conversational-volume speech, then set the Claude voice gain headroom once the badge clears. Deliver a sentence at the volume you would use when making a strong point, not at the level of polite background conversation. Also include deliberate plosive consonants: phrases like "pepperoni pizza" and "completely correct" generate the peak transients that most commonly trigger the detector. If the badge appears during this calibration run, reduce gain and repeat.

Setting a headroom buffer below the clipping ceiling

Once you confirm the ceiling gain level (highest gain where the badge does not appear during loudest speech), set your working gain 3–5 dB below that ceiling. Excited speech, unexpected questions that prompt a louder response, or the natural emphasis of explaining something technical all produce transients above your calibration baseline. The headroom buffer absorbs these excursions without triggering clipping. A 3 dB buffer is the minimum; 5 dB is recommended for conversational AI sessions where tone and volume vary unpredictably throughout the exchange.

When to use this

Check your microphone before any Claude Voice Mode session, particularly when switching between different input devices, audio setups, or environments. A 60-second pre-session check prevents transcription errors that slow the conversation.

Examples

USB condenser at high gain in small room

Before
Clipping badge appears during normal speech — Claude transcribes consonants incorrectly
After
Gain reduced 8 dB: no clipping badge, Noise Floor Grade Excellent — accurate transcription

Laptop built-in microphone in a coffee shop

Before
Noise floor −35 dBFS (Very Noisy) — multiple transcription errors per sentence
After
Switched to USB cardioid microphone: −58 dBFS — sentence-level accuracy restored
Sources
  1. 1.

    Anthropic, "Voice dictation," code.claude.com, accessed June 2026. https://code.claude.com/docs/en/voice-dictation

  2. 2.

    Wikipedia, "Clipping (audio)," en.wikipedia.org, accessed June 2026. https://en.wikipedia.org/wiki/Clipping_(audio)

  3. 3.

    IBM VoiceTIMES, "Audio Hardware Guidelines and Signal Specifications," public.dhe.ibm.com, August 1999. https://public.dhe.ibm.com/software/viavoicesdk/VoiceTIMES_HW_Spec.pdf

  4. 4.

    Alan Jongman, "Phonetics of Fricatives," kuppl.ku.edu, June 2024. https://kuppl.ku.edu/sites/kuppl/files/documents/publications/Jongman%20OREL%202024%20Phonetics%20of%20Fricatives.pdf

FAQ