Mic Test for AI Voice Assistants and Speech-to-Text

Check noise floor, clipping and latency before you talk to an AI voice assistant or record for Whisper, then fix the setting each assistant is most sensitive to.

Mic Test for AI Voice Assistants and Speech-to-Text

Every voice assistant listens through the same chain: your microphone, your operating system, then the assistant's own speech detection. When an assistant cuts you off, answers before you finish or keeps mishearing a word, the cause is usually something you can measure in that chain before you start talking. Three readings cover most of it: how loud the room is when you are silent, whether your loudest words hit the digital ceiling, and how much delay your audio path adds.

The first section sets those three targets once, because they apply to every assistant. The sections after it cover only what differs: how each assistant decides you have stopped talking, where its audio is processed and which setting it is most sensitive to.

Recommended tests for voice assistants and speech-to-text

  • Noise Floor Grade: background noise is what makes turn detection cut in early or wait too long
  • Clipping Detector: clipped consonants turn into misheard words that no server processing can repair
  • Clap Latency Test: extra delay on your side makes interruptions land late in live conversations

Opens the Microphone Quality, Noise & Latency Tester with this page's recommended tests marked.

Open in the tool →

The three readings every assistant needs

Aim for a Noise Floor Grade of Good or Excellent, which on this tool means a silent room reads below −50 dBFS. Assistants that detect speech by level, which most live voice modes do, compare your voice with the room. A Noisy grade (−50 to −40 dBFS) leaves less contrast, so a fan or a conversation next door can count as speech, and quiet word endings can count as silence. Fix the room before you touch any setting: closing a window, switching off a desk fan or moving the microphone closer to your mouth usually shifts the grade more than any software option.

Finding a safe gain window

Gain has two limits. Too low, and your voice sits close to the noise floor. Too high, and loud consonants clip at 0 dBFS, which flattens the waveform and adds distortion across the whole spectrum.1 Start with the Noise Floor Grade at your normal gain, then run the Clipping Detector for about 20 seconds while you speak at your loudest normal volume, including hard sounds like "p" and "b". If the badge appears, lower the gain a few steps and repeat. Then lower it a little more, because excited speech peaks higher than a calm test sentence.

Latency and Bluetooth

The Clap Latency Test grades the round trip through your browser and OS: under 60 ms is Good, and 60 to 100 ms is Noticeable. That delay matters for live assistants you can interrupt, because your interruption reaches the service later. Bluetooth headsets add the most. When a headset microphone is active, many headsets switch to the hands-free profile, whose narrowband mode carries telephone-quality audio, so the Frequency Response display drops off sharply above about 4 kHz.2 A wired or USB microphone avoids both problems.

Speaker bleed

Any assistant that talks back can hear itself. If its reply plays through speakers near an open microphone, the reply enters your input and can trigger a false interruption or a false wake word. Record an Echo Loopback while audio plays from your speakers to hear how much reaches the microphone, and use headphones when you can.

Headphones remove that path completely rather than reducing it, which is why they solve the problem on every assistant that speaks aloud. If you have to keep speakers, turn their volume down and move the microphone off the speaker axis, so the reply reaches the capsule at an angle its pickup pattern already rejects.

Sources
  1. 1.

    "Clipping (audio)," Wikipedia, accessed October 2026. https://en.wikipedia.org/wiki/Clipping_(audio)

  2. 2.

    Bluetooth SIG, "Headset/Hands-Free Profile (HFP)," bluetooth.com, accessed September 2026. https://www.bluetooth.com/specifications/specs/headset-hand-free-profile-1-7/

ChatGPT Voice Mode Mic Test

For ChatGPT Advanced Voice Mode, GPT-4o's audio model processes audio waveforms directly for end-to-end voice handling.1 Unlike previous OpenAI voice implementations that transcribed speech before generating a text response, Advanced Voice Mode keeps the audio path central to turn-taking and response timing. Background noise that a traditional speech-to-text engine might suppress through noise reduction degrades the interrupt detection mechanism, the part of the system that determines when you have stopped speaking and when to respond.

The noise floor threshold that matters for ChatGPT Advanced Voice Mode is approximately −50 dBFS or better. Above that level, background noise competes with the active speech signal in a way that confuses the end-of-speech detector, causing the model to respond prematurely or to miss the end of your sentence. Latency below 60ms round-trip (as measured by the Clap Latency Test) ensures that the model's audio output and your subsequent speech do not overlap with noticeable delay.2 Run the Noise Floor Grade, Clipping Detector, and Clap Latency Test as your pre-session checks before any Advanced Voice Mode conversation.

What to look for

  • noise floor above -45 dBFS degrades detection
  • 40 to 120 ms, often pushing total above 80 ms

Opens the Microphone Quality, Noise & Latency Tester with this section's reference values shown at the top of the tool.

Open in the tool →

Background noise and Advanced Voice Mode speech detection

Advanced Voice Mode processes audio in near real-time without a dedicated noise gate stage visible to the user. When background noise reaches the −40 dBFS range (the threshold the Noise Floor Grade labels Noisy), the model's speech activity detection activates for noise events as well as speech, producing false turn-taking signals. Furthermore, continuous background noise above −45 dBFS compresses the effective dynamic range available to GPT-4o's audio model for distinguishing your speech from silence.

Corrective actions that shift the grade fastest

Achieving a Noise Floor Grade of Good or Excellent before starting a voice session eliminates this problem reliably. If your initial reading is Noisy, take one corrective action, then retest before the session. Closing a window, turning off a nearby fan, or moving the microphone 10 cm closer to your mouth are the three fastest improvements. Each action alone can shift the grade by 5 to 10 dBFS. Running the test after each single change tells you which action had the most impact, so you know what to prioritize next time without guessing.

Target a Noise Floor Grade of Good or better, which corresponds to roughly −50 dBFS or lower, because Advanced Voice Mode's interrupt detection begins to misread background noise above that level. A single high-impact change that reaches the Good threshold is enough to clear a session, so you do not need every environmental fix in place before starting the conversation.

Interrupting ChatGPT mid-response

ChatGPT Advanced Voice Mode supports natural conversation with human-speed turn-taking, including the ability to interrupt the model mid-response.3 Round-trip latency above 80ms (the upper boundary of the Clap Latency Test's Good grade) adds a perceptible delay to the interrupt detection loop. Building on this, the model buffers audio on its end to manage network jitter, but the browser-to-network path latency adds on top of that server buffer. Keeping your Clap Latency Test result below 60ms eliminates the browser-side audio stack as a latency contribution.4

When you interrupt Advanced Voice Mode, the interrupt signal must reach OpenAI's servers before the next audio frame is generated and streamed back to your browser. If your browser-side latency is 80ms or higher, the server has already committed to additional output frames by the time your interrupt arrives, which means the model continues speaking for a noticeable window before it registers that you have started talking over it. Reducing browser-side latency tightens this window and makes the conversational back-and-forth feel closer to a natural phone call where both participants can react instantly.

Clipping and the speech band GPT-4o listens to

GPT-4o's audio model processes speech in the 80 Hz–8 kHz fundamental range where human voice energy is concentrated.5 Clipping at 0 dBFS introduces harmonic distortion that extends far above this range, creating an artificial spectral signature that the model processes less effectively. Consequently, the Clipping Detector is particularly important here: even brief clipping during peak consonant sounds corrupts the audio frame the model uses to parse phonemes.6 Aim for a Frequency Response display that shows consistent energy through the speech range without sharp peaks indicating gain staging problems.

Unlike traditional pipelines that separate transcription from language processing, GPT-4o's audio model maps raw waveforms directly to semantic representations without an intermediate text stage. This means clipped samples do not just produce a wrong letter in a transcript; they distort the entire audio frame the model uses to determine what was said at that moment. The harmonic artifacts from clipping spread across all frequencies simultaneously, confusing the model's attention mechanism in ways that are fundamentally harder to recover from than a simple noise-floor issue that spectral subtraction could address.

Gain staging for Advanced Voice Mode sessions

Before an Advanced Voice Mode session, the gain staging window that satisfies both the Noise Floor Grade and the Clipping Detector simultaneously defines your safe operating range. For most microphones in a quiet room, this window spans 15–20 dB of gain adjustment: below the lower bound, the Noise Floor Grade is Noisy because the signal is too weak relative to ambient noise; above the upper bound, the Clipping Detector fires during peak consonant sounds. Finding this window before the session prevents the most common cause of recognition errors: gain set too high with frequent clipping, or too low with a poor speech-to-noise ratio.

Using both detectors as a gain staging bracket

Run the Noise Floor Grade at your starting gain and confirm it is Good or Excellent. Then run the Clipping Detector while speaking at your loudest normal voice for 30 seconds, including deliberate plosive consonants. If no badge appears, you are inside the safe window. If the badge appears, reduce gain by 3 dB and repeat the Clipping Detector test. This dual-test approach brackets the gain from both ends and confirms the safe operating point with a single adjustment sequence rather than separate gain searches for each constraint independently.

Reading the Frequency Response display before Advanced Voice Mode

The Frequency Response display shows whether your microphone is delivering speech content in the 80 Hz–8 kHz range that GPT-4o's audio model processes. Speak a sustained vowel while watching the display. Energy should appear consistently across the 200 Hz–4 kHz fundamental speech range. If the display shows a steep drop-off above 4 kHz, your microphone may be in HFP Bluetooth mode or using a narrow-bandwidth codec; switch to a wired connection.7

Identifying room problems in the display before a session

The Frequency Response display also reveals room problems that are not visible in the single-number Noise Floor Grade, which is why you should read the ChatGPT voice Frequency Response before every session. A consistent energy spike between 100–200 Hz indicates HVAC noise entering the measurement; a series of regularly spaced dips across the midrange suggests comb filtering from a nearby reflective surface. Either of these may not push the Noise Floor Grade into the Noisy range alone, but they reduce the speech signal quality available to Advanced Voice Mode's audio processing at specific frequency bands where phoneme discrimination depends on consistent spectral content.

When to use this

Run this check before any GPT-4o Advanced Voice Mode conversation, especially in noisy environments, or when you notice the model cutting you off mid-sentence or missing the end of your thoughts.

Examples

Open-plan office with background conversation

Before
Noise floor at −38 dBFS (Noisy) — model interrupts prematurely during pauses in speech
After
Moved to a quiet conference room: noise floor −62 dBFS (Excellent) — model waits for full sentences

Home desk with nearby AC unit

Before
Noise floor −44 dBFS, Clap Latency 75ms — occasional missed turn signals and hesitation
After
AC switched off, USB microphone moved to direct port: −59 dBFS noise floor, 42ms latency — clean conversation
Sources
  1. 1.

    OpenAI, "Hello GPT-4o," openai.com, May 2024. https://openai.com/index/hello-gpt-4o/

  2. 2.

    OpenAI and LiveKit, "OpenAI and LiveKit partner to turn Advanced Voice into an API," livekit.com, March 2024. https://livekit.com/blog/openai-livekit-partnership-advanced-voice-realtime-api

  3. 3.

    OpenAI, "Voice Activity Detection (VAD)," developers.openai.com, accessed June 2026. https://developers.openai.com/api/docs/guides/realtime-vad

  4. 4.

    Alexis Conneau / TechCrunch, "The creator of ChatGPT's voice wants to build the tech from 'Her,'" techcrunch.com, December 2024. https://techcrunch.com/2024/12/09/the-creator-of-chatgpts-voice-wants-to-build-the-tech-from-her-minus-the-dystopia/

  5. 5.

    Ars Technica, "Major ChatGPT-4o update allows audio-video talks with emotional chatbot," arstechnica.com, May 2024. https://arstechnica.com/information-technology/2024/05/chatgpt-4o-lets-you-have-real-time-audio-video-conversations-with-emotional-chatbot/

  6. 6.

    Wikipedia, "Formant," en.wikipedia.org, accessed June 2026. https://en.wikipedia.org/wiki/Formant

  7. 7.

    Apple Siri Team, "Improving Neural Network Acoustic Models by Cross-bandwidth Initialization," machinelearning.apple.com, August 2017. https://machinelearning.apple.com/research/cross-initialization

FAQ

The API processes the raw browser audio stream. Some noise reduction may occur on OpenAI's servers, but this is not documented behavior and cannot be relied upon to compensate for a poor noise floor on your end.

End-of-speech detection identifies silence following speech. When your noise floor is above −45 dBFS, CapyToolkit's Noise Floor Grade shows the background is not quiet enough for the detector to reliably identify silence. It may wait indefinitely for a deeper pause that never arrives.

A cardioid condenser or a close-miked dynamic both work well. Dynamic microphones outperform condensers in noisy environments because their reduced sensitivity to ambient sound produces better Noise Floor Grades without soundproofing.

Yes. Traditional transcription engines can partially recover clipped audio through their preprocessing stages. Advanced Voice Mode processes waveforms directly, so clipping artifacts affect audio model inference more immediately.

Bluetooth latency typically adds 40 to 120ms on top of browser-side processing, often pushing the total Clap Latency Test result above 80ms. Wired USB or 3.5mm connections produce more reliable latency grades.

Claude Voice Mode Mic Test

Claude Voice Mode depends on a speech-to-text stage before response generation. Unlike end-to-end audio models, this pipeline has distinct stages: audio capture, speech recognition, and language model inference.1 Clipping at the audio capture stage introduces distortion that corrupts the speech recognition stage specifically: the waveform shape the recognition model depends on is destroyed by digital saturation, producing transcript errors that the language model then tries to interpret from damaged input.2

A noise floor below −50 dBFS is the threshold where speech recognition errors from background noise become negligible in Claude's voice pipeline for standard accents and speaking rates. Above that level, error rates in the transcription stage increase gradually, though they rarely become catastrophic until the floor exceeds −40 dBFS. The Noise Floor Grade, Clipping Detector, and Frequency Response display are the three tests to run before a voice session. Specifically, ensure the Clipping Detector shows no badge during your loudest normal speech, and that the Noise Floor Grade reads Good or Excellent.

What to look for

  • above -40 dBFS noise floor

Opens the Microphone Quality, Noise & Latency Tester with this section's reference values shown at the top of the tool.

Open in the tool →

Room noise before server-side transcription

Claude's voice pipeline transcribes audio before the language model processes it, and that transcription happens on Anthropic's servers rather than on your machine: the docs state that recorded audio is streamed to Anthropic for transcription and is never processed locally.3 Background noise that sits between −50 dBFS and −40 dBFS (the Noisy range on the Noise Floor Grade) reduces transcription accuracy by competing with the speech signal in critical frequency bands. Consonants, particularly the sibilant frequencies between 3–8 kHz, are most affected because their energy is weaker relative to fundamental vowel frequencies. Furthermore, sustained background noise causes the speech recognition model to allocate computational attention to modeling noise rather than speech, which reduces confidence scores on phoneme-to-word mapping.

Sibilant sounds like "s," "f," and "sh" produce their distinguishing acoustic energy in the 3 to 8 kHz range, where amplitude is naturally 15 to 25 dB below the fundamental vowel energy concentrated below 1 kHz.4 When background noise fills this upper frequency band, the speech recognition model receives a signal where the noise floor and the sibilant energy occupy the same amplitude range, making it impossible to distinguish between a genuine "s" sound and a noise burst that happens to have similar spectral character. Running the Noise Floor Grade before each session catches this problem before it manifests as repeated transcription errors on the most common consonants in English.

Turn-taking with silence detection and push-to-talk

Claude Voice Mode responds after your speech turn ends, using silence detection to identify turn boundaries. Claude Code's dictation mode makes the same assumption explicitly: recording runs while the push-to-talk key is held and stops when you release it, which is a fixed turn boundary rather than an inferred one.5 Round-trip latency measured by the Clap Latency Test affects the delay you perceive between asking a question and hearing Claude begin its response. Latency below 60ms keeps the browser-side audio stack out of the critical path; at that level, network and server processing time dominate the perceived delay. Building on this, if the Echo Loopback test reveals significant distortion in the playback, investigate whether your monitoring setup is creating acoustic feedback into the microphone during Claude's audio output.

Claude's turn-taking depends on detecting a sufficient drop in audio amplitude after you finish speaking, which signals that the model should begin generating its response. When browser-side latency is high, the silence detection window at the server receives your audio delayed by the full round-trip time, which means the model waits longer than necessary before responding and the conversation develops an unnatural rhythm where each turn begins with a perceptible pause. Keeping the Clap Latency Test result in the Good range or better minimizes this browser-side contribution and lets the server-side silence detection operate on the most current audio available.

Clipped plosives and transcription errors

Clipping produces harmonic distortion that extends across the entire spectrum above the clipped frequency.2 For Claude's speech recognition pipeline, clipping during peak speech levels, particularly during plosive consonants like p and b, corrupts the short audio frames the acoustic model processes. The Clipping Detector triggers when more than 1% of samples in a frame hit 0 dBFS. Even brief episodes at this level during normal speech indicate gain staging that is too high. The Frequency Response display helps confirm that the gain setting that eliminates clipping still delivers sufficient energy through the 200 Hz–4 kHz fundamental speech range.

The language model stage that follows transcription can often infer the correct word from surrounding context even when the transcript contains minor errors, but the speech recognition stage that produces the initial transcript has no such contextual safety net for individual phonemes. A clipped "p" sound at the start of a word produces a distorted waveform that the acoustic model maps to an entirely wrong phoneme candidate, and the language model downstream receives this wrong phoneme sequence with no way to know the original signal was corrupted by digital saturation rather than being a genuine speech sound. This is why the Clipping Detector matters more for Claude Voice Mode than for text-based interactions.

Gain staging for Claude's speech recognition pipeline

Because Claude's voice pipeline uses a speech recognition stage before the language model processes your input, gain staging affects transcript accuracy directly rather than indirectly through a processing buffer. Starting at 60% OS input gain provides enough signal for a usable Noise Floor Grade in most quiet rooms while leaving headroom below the Clipping Detector threshold for a typical voice. Run the Noise Floor Grade first; if the result is Noisy, increase gain by 5% and repeat. If the result is Good or Excellent, run the Clipping Detector while speaking at maximum normal volume before beginning any session.

Finding the optimal gain bracket in under five steps

The search for the optimal gain setting takes five adjustment steps at most in a typical room. Start at 40% OS input gain, run the Noise Floor Grade, and note the result. Increase to 50% and repeat. At each step, the grade improves until room noise becomes the dominant source and further gain increases stop improving it. That inflection point is the optimal gain setting. Confirm by running the Clipping Detector at that gain; if no badge appears during loud speech, the setting is both noise-floor-optimal and clipping-safe.

Confirming the safe window before each session

After finding the gain level where the Noise Floor Grade is Good and the Clipping Detector shows no badge during loud speech, document the OS gain slider position with a screenshot. Revisit this calibration when you change rooms, change hardware, or reconnect the microphone after a period of non-use. Gain staging that was optimal in a previous session may be too high if the OS default communications device changed and selected a different microphone at its default gain setting. The two-minute dual-test sequence prevents this common source of degraded voice sessions.

The dual-test sequence is the Noise Floor Grade followed immediately by the Clipping Detector at the same gain, completed in under two minutes. Keeping the screenshot of the OS slider beside the microphone means you can restore the exact position after any device change without re-running the full search, which is the practical value of documenting the calibration rather than relying on memory of a percentage.

Understanding how soft speech affects clipping risk

Clipping risk in Claude Voice Mode is determined by your loudest typical speech, not your average or softest speech. The Clipping Detector threshold of 1% samples at 0 dBFS means it fires when any speech peak exceeds the digital ceiling, regardless of how quiet the rest of the recording is. If your gain staging is calibrated to your whispered or quiet voice, the gain may be set too high to handle emphatic speech, technical explanations delivered with more force, or the natural loudness increase that occurs when speaking about something important.

Calibrating with loud consonants as the test signal

During the Clipping Detector calibration run, include deliberate emphatic speech rather than only conversational-volume speech, then set the Claude voice gain headroom once the badge clears. Deliver a sentence at the volume you would use when making a strong point, not at the level of polite background conversation. Also include deliberate plosive consonants: phrases like "pepperoni pizza" and "completely correct" generate the peak transients that most commonly trigger the detector. If the badge appears during this calibration run, reduce gain and repeat.

Setting a headroom buffer below the clipping ceiling

Once you confirm the ceiling gain level (highest gain where the badge does not appear during loudest speech), set your working gain 3–5 dB below that ceiling. Excited speech, unexpected questions that prompt a louder response, or the natural emphasis of explaining something technical all produce transients above your calibration baseline. The headroom buffer absorbs these excursions without triggering clipping. A 3 dB buffer is the minimum; 5 dB is recommended for conversational AI sessions where tone and volume vary unpredictably throughout the exchange.

When to use this

Check your microphone before any Claude Voice Mode session, particularly when switching between different input devices, audio setups, or environments. A 60-second pre-session check prevents transcription errors that slow the conversation.

Examples

USB condenser at high gain in small room

Before
Clipping badge appears during normal speech — Claude transcribes consonants incorrectly
After
Gain reduced 8 dB: no clipping badge, Noise Floor Grade Excellent — accurate transcription

Laptop built-in microphone in a coffee shop

Before
Noise floor −35 dBFS (Very Noisy) — multiple transcription errors per sentence
After
Switched to USB cardioid microphone: −58 dBFS — sentence-level accuracy restored
Sources
  1. 1.

    Anthropic, "Voice dictation," code.claude.com, accessed June 2026. https://code.claude.com/docs/en/voice-dictation

  2. 2.

    Wikipedia, "Clipping (audio)," en.wikipedia.org, accessed June 2026. https://en.wikipedia.org/wiki/Clipping_(audio)

  3. 3.

    IBM VoiceTIMES, "Audio Hardware Guidelines and Signal Specifications," public.dhe.ibm.com, August 1999. https://public.dhe.ibm.com/software/viavoicesdk/VoiceTIMES_HW_Spec.pdf

  4. 4.

    Alan Jongman, "Phonetics of Fricatives," kuppl.ku.edu, June 2024. https://kuppl.ku.edu/sites/kuppl/files/documents/publications/Jongman%20OREL%202024%20Phonetics%20of%20Fricatives.pdf

  5. 5.

    Anthropic, "Voice dictation," code.claude.com, accessed September 2026. https://code.claude.com/docs/en/voice-dictation#requirements

FAQ

Transcription errors increase, particularly for proper nouns and technical terms. CapyToolkit's Noise Floor Grade helps you catch this before the session. Claude's language model can often infer context from surrounding words, but accuracy decreases noticeably above −40 dBFS noise floor.

Transcription models handle accents better at lower noise floors. In the Noisy range, background noise competes with the lower-amplitude speech features that distinguish accented phonemes from their standard counterparts.

Yes. If speaker audio bleeds back into the microphone during Claude's response, it appears as input in your next turn. Run the Echo Loopback test to verify how much playback audio your microphone captures.

Voice activity detection works well when the Noise Floor Grade is Good or Excellent. In the Noisy range, push-to-talk is more reliable because it eliminates false activations from background noise.

Claude's voice pipeline accepts standard browser audio streams, typically 16-bit or 24-bit depending on the microphone and OS. The Noise Floor Grade measures performance at whatever bit depth your microphone delivers.

Gemini Live Mic Test

In Gemini Live, speech moves through a streaming audio pipeline with automatic gain control applied on Google's servers. AGC normalizes your input volume before the speech recognition stage1, which means moderate noise floors that would cause errors in a fixed-gain pipeline are handled more robustly. Yet AGC has limits: when the noise floor exceeds −35 dBFS (the Very Noisy range on the Noise Floor Grade), gain normalization amplifies background noise proportionally, pushing it into the same amplitude range as speech and causing the model to interpret noise as speech activity.

The Frequency Response display reveals an important consideration for Gemini Live: the service processes audio optimized for the 300 Hz–3.4 kHz telephone bandwidth range2, though wider input is accepted. Microphones that roll off below 200 Hz or above 8 kHz still work effectively, making noise floor and clipping performance more important than frequency response width. The Clap Latency Test is particularly relevant for Gemini Live's conversational turn-taking: the service uses interruption-capable processing, and high browser-side latency creates perceptible hesitation in the conversation flow.

What to look for

  • 40 to 120 ms

Opens the Microphone Quality, Noise & Latency Tester with this section's reference values shown at the top of the tool.

Open in the tool →

Gemini Live gain control in a noisy room

Gemini Live's AGC normalizes input amplitude but cannot separate speech from noise when both occupy similar frequency bands and amplitude ranges. At noise floor levels between −45 dBFS and −35 dBFS, AGC amplifies the combination of speech and noise together, reducing the speech-to-noise ratio at the point of recognition rather than improving it3. Furthermore, sustained noise sources that repeat rhythmically (fan blades, air conditioning cycles) can confuse the speech activity detector, causing the service to begin processing and then abort mid-response. A Noise Floor Grade of Good or Excellent eliminates this risk reliably. The key threshold to remember is that AGC becomes counterproductive when the noise floor exceeds −45 dBFS, because at that point the gain normalization stage amplifies background noise proportionally with your voice rather than isolating the speech signal, which degrades the speech-to-noise ratio that the recognition model depends on for accurate transcription.

Why AGC cannot fix a poor noise floor

Automatic gain control on Google's servers adjusts the overall amplitude of the incoming audio stream to normalize volume levels across different microphones and distances, but it processes the combined signal of speech and noise together without the ability to distinguish between the two. When background noise occupies the same frequency bands as your voice, raising the gain to make speech louder also makes the noise louder by the same amount, which means the speech-to-noise ratio at the recognition stage remains unchanged despite the AGC adjustment.

This is why the practical target remains a Noise Floor Grade of Good or better before any Gemini Live session, because once the floor crosses the −45 dBFS point the AGC stage begins scaling noise upward alongside your voice rather than isolating it. A clean input delivered to the server gives the recognition model a stable speech-to-noise ratio that no amount of server-side normalization can reconstruct after the fact.

Interrupting Gemini Live while it answers

Gemini Live supports interruption, meaning you can speak over the model's response to redirect the conversation. This feature works best when browser-side audio latency is below 80ms. At higher latency, the interrupt signal reaches the server after additional model output has already been generated, making clean interruptions less predictable. The Clap Latency Test measures the full browser audio round-trip. Keeping this below 60ms gives Gemini Live's interruption handling the tightest possible timing window from your side of the connection.

When you speak over Gemini Live's output to change the subject or correct a misunderstanding, your interrupt audio must travel through the browser's WebAudio pipeline, across the network, and into the server's streaming inference loop before the model registers that a new speaker has taken control. Every millisecond of browser-side latency adds directly to this chain, which means a Clap Latency Test result of 90ms creates a noticeably longer window where the model continues generating its original response before your redirect arrives and gets processed.

Clipping that happens before Gemini hears you

AGC on the server side cannot recover samples that were clipped before transmission. Once the waveform is saturated at 0 dBFS in the browser's audio stream, the sample value is corrupted before leaving your device4. The Clipping Detector running in the browser is therefore the correct place to catch this: it flags clipping before the signal reaches Gemini Live's servers. Frequency response outside the core speech range matters less for Gemini Live than for wideband audio AI services, because the recognition model targets the 300 Hz–3.4 kHz band most heavily.

Why clipped samples bypass AGC recovery entirely

Server-side AGC operates on the digital samples it receives, adjusting their amplitude up or down to normalize the overall level. But when a sample is already pinned at 0 dBFS because of clipping in the browser's audio pipeline, the AGC stage has no information about what the original waveform peak looked like before it was truncated. The saturated sample value is the only data available, so the AGC can only scale a corrupted value that no longer represents the original sound pressure wave. This is why the Clipping Detector in the browser catches the problem at the only point where the original signal is still intact and the damage can actually be prevented.

Gemini Live's frequency range and microphone selection

Comparing microphone types for Gemini Live reveals that frequency response width matters less than noise floor and clipping performance. The service processes audio most effectively in the 300 Hz–3.4 kHz telephone bandwidth range, where its speech recognition model is most heavily weighted. A dynamic microphone with a frequency response that rolls off above 12 kHz still covers this band fully, while its lower sensitivity compared to a condenser produces a better Noise Floor Grade in noisy environments because it rejects ambient room sound more effectively at equivalent gain settings.

At equivalent OS gain settings, a dynamic microphone produces a quieter Noise Floor Grade than a condenser in the same room because its capsule is less sensitive to ambient sound5. The useful comparison is not at identical gain settings but at gain-matched conditions: set each microphone to the OS gain level that achieves an identical Noise Floor Grade result. Condenser microphones may require substantially lower gain to match a dynamic microphone's noise floor in a live room.

Broadcast-style dynamic microphones typically produce Noise Floor Grades of Excellent in the same room where a budget condenser might show Noisy, because the dynamic's lower sensitivity requires gain levels where room noise contributes less to the measurement. For Gemini Live sessions in an office environment with background noise, a dynamic microphone often outperforms a condenser despite the condenser's theoretically better specifications: the lower sensitivity advantage overcomes the condenser's lower self-noise when room noise is the dominant floor source.

Using multiple Noise Floor Grade runs to identify AGC instability

Gemini Live's server-side AGC can produce inconsistent behavior when your noise floor sits close to the edge of what AGC handles cleanly. If your Noise Floor Grade shows more than 6–8 dBFS variation across five consecutive runs in unchanged conditions, OS-level Automatic Gain Control may be adjusting the input amplification between runs6, which produces exactly this kind of run-to-run variation and causes the service to receive audio at fluctuating amplitudes.

Identifying AGC behavior from test variation

Run the Noise Floor Grade five times without changing any settings or room conditions. If the first two runs return Good and subsequent runs return Excellent, the room just needed to settle. If the variation is persistent and unpredictable across all five runs, spot Gemini Live AGC instability and check whether Windows AGC is enabled in your audio device properties. Windows AGC adjusts input gain dynamically and creates this pattern of test-to-test variation. Disabling Windows AGC and rerunning five tests confirms whether the variation disappears: if it does, OS AGC was the source, not the room or the server's processing.

When to use this

Use this check before Gemini Live conversations when you are in an environment with variable background noise, especially when AGC behavior feels inconsistent or when the service misinterprets background sounds as speech.

Examples

Home office with ceiling fan running

Before
Noise floor −40 dBFS — AGC amplifies fan noise to speech level, causing false activations
After
Fan switched off: noise floor −64 dBFS (Excellent) — clean conversation without false starts

Condenser mic at high OS gain

Before
Clipping badge triggers on loud consonants — Gemini Live cuts off and restarts recognition
After
OS gain reduced 12 dB: no clipping, Excellent noise floor — uninterrupted sessions
Sources
  1. 1.

    Google, "Live API capabilities guide," ai.google.dev, accessed June 2026. https://ai.google.dev/gemini-api/docs/live-api/capabilities

  2. 2.

    Wikipedia, "Wideband audio," en.wikipedia.org, accessed June 2026. https://en.wikipedia.org/wiki/Wideband_audio

  3. 3.

    WebRTC, "AGC2 Common Constants," chromium.googlesource.com, accessed June 2026. https://webrtc.googlesource.com/src/+/87b86acde990a0288b2c75be4f03d8bd5e1be74b/modules/audio_processing/agc2/agc2_common.h

  4. 4.

    MDN, "Web Audio API," developer.mozilla.org, accessed June 2026. https://developer.mozilla.org/en-US/docs/Web/API/Web_Audio_API

  5. 5.

    Neumann, "What is Sensitivity?," neumann.com, accessed June 2026. https://www.neumann.com/en-us/knowledge-base/neumann-im-homestudio/homestudio-academy/what-is-sensitivity

  6. 6.

    Microsoft, "KSPROPERTY_AUDIO_DEV_SPECIFIC," learn.microsoft.com, accessed June 2026. https://learn.microsoft.com/en-us/windows-hardware/drivers/audio/ksproperty-audio-dev-specific

FAQ

AGC normalizes volume but cannot fix a poor signal-to-noise ratio. CapyToolkit's Noise Floor Grade catches that problem before AGC: if your noise floor is in the Noisy range, AGC amplifies noise proportionally with speech. A clean input before AGC always produces better results.

Run the Frequency Response display while speaking normally. If you see very low energy through 300 Hz to 2 kHz, your microphone may be positioned too far from your mouth or your OS gain may be too low.

The recognition engine is the same across tiers. Audio quality improvement from a better microphone benefits all tiers equally. The model does not apply additional processing for paid users.

Bluetooth adds 40 to 120ms to browser-side latency, which typically pushes the Clap Latency Test result into the Noticeable or High range. The conversation still functions but interruptions feel slightly delayed.

Gemini Live uses a different model pipeline from Google Assistant. The test results and preparation steps are similar, but the recognition and response behavior comes from the Gemini model family.

Microsoft Copilot Voice Mic Test

On Windows, Microsoft Copilot Voice operates through the Windows audio stack by design, not as a workaround but as an intentional integration. Every audio enhancement in the Windows system applies before Copilot Voice processes input: Noise Suppression, Acoustic Echo Cancellation, and Automatic Gain Control in the Windows audio device properties panel all modify the signal before Copilot Voice receives it1. Consequently, the Noise Floor Grade and Clipping Detector in this browser-based tool measure the post-enhancement signal that Copilot Voice actually sees.

Gain staging set in the Windows sound control panel is the primary dial for the Noise Floor Grade and Clipping Detector results. The Windows microphone level slider controls input amplification at the OS level, and this applies before any application-specific processing. Because both Copilot Voice and this testing tool read from the same OS audio stream, calibrating your Windows microphone gain with the Noise Floor Grade test directly calibrates the input level that Copilot Voice receives2. The Clap Latency Test result reflects the Windows audio buffer settings, which are configurable via audio device properties.

What to look for

  • Noisy around -40 dBFS, Excellent around -60 dBFS
  • under 80 ms round trip
  • reduces buffering by 15-40 ms in most configurations

Run the Noise Floor Grade with Windows enhancements disabled first for a clean hardware baseline, then compare with enhancements enabled.

Opens the Microphone Quality, Noise & Latency Tester with this section's reference values shown at the top of the tool.

Open in the tool →

Windows noise suppression and the measured floor

Windows Noise Suppression, when enabled, can dramatically reduce the measured noise floor, sometimes from Noisy (−40 dBFS) to Excellent (−60 dBFS). Yet this apparent improvement may disguise rather than solve an underlying problem: Windows classes noise suppression and automatic gain control as adaptive processing that its raw capture mode must leave out, so a reading taken with them on describes the processed signal rather than the microphone itself3.

Hardware baseline versus enhanced measurement

Run the Noise Floor Grade with Windows enhancements disabled first to get the hardware baseline, then compare with enhancements enabled. Copilot Voice performs best when the hardware input is already clean rather than relying on suppression processing. The difference between the two readings reveals how much Windows audio processing is contributing to your noise floor result. If the hardware baseline is already Good or Excellent, Windows enhancements are unnecessary and may introduce processing artifacts that degrade voice quality. If the hardware baseline is Noisy, enhancements can help, but physical room treatment produces a cleaner result without the warbling or metallic side effects that suppression algorithms sometimes introduce.

Windows audio buffers and Copilot response delay

Windows audio buffer settings determine the primary source of browser-side audio latency in this test. Larger audio buffers reduce CPU interrupts at the cost of higher latency. The Clap Latency Test shows the total round-trip, which includes Windows audio processing, the browser's WebAudio implementation, and network overhead. For Copilot Voice specifically, keeping latency below 80ms ensures the conversation turn-taking feels responsive. Building on this, Windows Exclusive Mode in audio device advanced settings allows applications to bypass the shared audio session and access the hardware directly, reducing latency by 20 to 40ms in some configurations4.

How Windows audio buffers stack across the pipeline

The total round-trip latency measured by the Clap Latency Test is the sum of multiple independent buffer stages: the Windows audio driver buffer, the Windows mixer buffer if shared mode is active, and the browser's WebAudio internal buffer. Each stage adds its own delay independently, so a system with three 10ms buffers produces 30ms of audio delay before the signal even reaches the network stack. Understanding this stacking behavior helps you prioritize which buffer to shrink first when the Clap Latency Test returns a result above the Good threshold.

Windows echo cancellation and clipped speech

Windows Acoustic Echo Cancellation, when active, processes both the microphone input and speaker output simultaneously. This helps prevent speaker audio from feeding back into the microphone, which matters for the Echo Loopback test. However, echo cancellation applies a nonlinear processing stage that can significantly alter the Frequency Response display in ways that do not reflect your microphone's actual hardware characteristics, because the algorithm continuously models and subtracts the estimated speaker contribution from the microphone signal5.

Consequently, the Frequency Response reading on Windows with echo cancellation enabled may not represent your microphone's hardware response accurately; it reflects the post-processing output, which is also what Copilot Voice receives. This distinction matters when diagnosing audio quality problems: if the Frequency Response display looks uneven or shows unexpected rolloff with enhancements enabled, the issue may be in the Windows processing layer rather than the microphone hardware. Disable all enhancements and compare the display again. If the response smooths out, Windows processing was the cause, and you can selectively re-enable individual enhancements to identify which one introduces the artifact.

Diagnosing Windows audio enhancement conflicts

In Windows audio device settings, four common enhancements can affect the Noise Floor Grade in opposite directions: Noise Suppression typically improves it, Automatic Gain Control destabilizes it across runs, and Bass Boost raises the low-frequency noise contribution. Acoustic Echo Cancellation has minimal effect on a measurement taken in silence. Disable all enhancements first and run the Noise Floor Grade for a clean hardware baseline. If the baseline is Good or Excellent, enhancements are unnecessary for your environment. If the baseline is Noisy, re-enable Noise Suppression only and run the test again to see whether that single enhancement produces enough improvement to bring the grade into the Good range without introducing processing artifacts.

If Noise Suppression alone improves the grade by 5 or more dBFS without introducing audible artifacts in the Echo Loopback playback, keep it enabled. If Noise Suppression improves the grade but creates warbling or metallic quality in the Echo Loopback, disable it; the hardware noise floor is too high for clean suppression processing, and physical environment improvement is required instead. Automatic Gain Control changes the Noise Floor Grade result unpredictably across runs because it adjusts gain dynamically6; always disable AGC before running the Noise Floor Grade to get a stable and repeatable measurement.

Re-enabling all four enhancements simultaneously after a Noisy baseline makes it impossible to identify which specific processing stage is responsible for the improvement or for any new artifacts that appear. By enabling one enhancement at a time and running the Noise Floor Grade after each change, you build a clear picture of what each stage contributes: Noise Suppression might improve the grade by 8 dBFS while adding slight warbling, while Bass Boost might worsen it by 3 dBFS with no audible benefit. This systematic single-variable approach takes slightly longer than testing all combinations but produces a definitive answer about which enhancements are genuinely helping your specific microphone and room combination.

Enabling Exclusive Mode to reduce Copilot Voice latency

Exclusive Mode in Windows audio device settings allows a single application to take direct control of the audio hardware, bypassing the shared audio mixer and reducing buffering by 15–40ms in most configurations. Copilot Voice on Windows benefits from this: the shared audio session adds processing stages that accumulate latency. Enabling Exclusive Mode requires that no other application is simultaneously using the microphone; only one application can hold exclusive access at a time, so the improvement applies only when no other audio application has opened the device.

In Windows Sound Control Panel, right-click your microphone, select Properties, and navigate to the Advanced tab. Check both "Allow applications to take exclusive control of this device" and "Give exclusive mode applications priority." Close the dialog, then run the Clap Latency Test and compare the result against your baseline. A reduction of 15–40ms confirms Exclusive Mode is active and effective for your setup. If the Clap Latency Test shows no improvement, another buffer stage in the pipeline is the dominant contributor to latency rather than the shared audio session overhead.

When Exclusive Mode conflicts with other applications

Exclusive Mode creates a compatibility constraint: any application that tries to open the microphone while Copilot Voice holds exclusive access will fail silently or fall back to a lower-quality shared mode. Browser-based tools including this mic test page cannot open the device while an exclusive application holds it. Close Copilot Voice before running any diagnostic test here, then re-open it after confirming your baseline. This sequence ensures both the test and Copilot Voice access the microphone at its configured sample rate and bit depth rather than through a resampled shared session.

Treat the close-and-reopen step as mandatory rather than optional, because a shared session that resamples the device to a different rate than the microphone native rate adds both latency and a slight quality loss that the test would otherwise attribute to the hardware. Keeping Copilot Voice closed while you confirm Copilot Voice Exclusive Mode latency guarantees the Clap Latency Test and Noise Floor Grade reflect the true raw device rather than a software-mediated stream.

When to use this

Use this check before any Copilot Voice session on Windows, particularly after updating Windows audio drivers, changing microphone devices, or modifying Windows audio enhancement settings.

Examples

Windows 11 with all audio enhancements enabled

Before
Noise Floor Grade reads Excellent but speech sounds warbling — suppression artifacts in processed signal
After
Disabled noise suppression: hardware noise floor −55 dBFS (Good) — natural voice quality without artifacts

USB microphone through Windows Exclusive Mode

Before
Clap Latency Test: 85ms — slightly above Good threshold
After
Enabled Exclusive Mode via advanced properties: 48ms — Good grade
Sources
  1. 1.

    Microsoft, "Audio Processing Object Architecture," learn.microsoft.com, accessed June 2026. https://learn.microsoft.com/en-us/windows-hardware/drivers/audio/audio-processing-object-architecture

  2. 2.

    Microsoft, "IAudioClient::Initialize," learn.microsoft.com, accessed June 2026. https://learn.microsoft.com/en-us/windows/win32/api/audioclient/nf-audioclient-iaudioclient-initialize

  3. 3.

    Microsoft, "Audio Signal Processing Modes," learn.microsoft.com, accessed October 2026. https://learn.microsoft.com/en-us/windows-hardware/drivers/audio/audio-signal-processing-modes

  4. 4.

    Microsoft, "Fix distorted or crackling audio in Windows," support.microsoft.com, accessed June 2026. https://support.microsoft.com/en-us/windows/fix-distorted-or-crackling-audio-in-windows-5304e452-38a6-4f3b-83cd-664beb3e68aa

  5. 5.

    Wikipedia, "Automatic gain control," en.wikipedia.org, accessed June 2026. https://en.wikipedia.org/wiki/Automatic_Gain_Control

  6. 6.

    Microsoft, "Fix microphone problems," support.microsoft.com, accessed June 2026. https://support.microsoft.com/en-us/windows/fix-microphone-problems-5f230348-106d-bfa4-1db5-336f35576011

FAQ

Test both ways with the Noise Floor Grade. If enhancements move the grade from Noisy to Good without audible artifacts in the Echo Loopback test, keep them enabled. If you hear warbling or suppression artifacts, disable enhancements and improve the physical environment instead.

Copilot Voice uses Microsoft's Azure Speech pipeline, which is separate from Windows built-in speech recognition. Windows audio enhancements still apply to the input stream before it reaches either system.

Windows updates sometimes reset audio driver settings to default. CapyToolkit's Clap Latency Test can reveal the change: check your audio device advanced properties and verify the sample rate matches your microphone's native rate. Mismatched rates cause the OS to resample, adding latency.

If microphone access is denied at the Windows privacy settings level, no application, including this browser tool, can access audio input. Run the Noise Floor Grade first to confirm your microphone is accessible before troubleshooting Copilot Voice.

Copilot Voice on mobile runs through the iOS or Android audio stack rather than Windows. The noise floor and latency requirements are similar, but the test tool on a mobile browser uses the same getUserMedia path as desktop.

OpenAI Whisper Mic Test

For transcription work, OpenAI Whisper is a speech recognition model available via API that accepts audio file uploads for batch transcription. Unlike real-time voice AI services, Whisper processes audio after the fact rather than in a continuous stream. This distinction matters for which tests apply: latency is irrelevant to batch transcription, but noise floor accuracy directly predicts how many corrections you will need to make to Whisper's output.

OpenAI trained Whisper on 680,000 hours of web audio and reports that the scale of that data improved its robustness to accents, background noise and technical language1. Yet the noise floor remains the strongest single predictor of transcription accuracy even for Whisper: a Noise Floor Grade of Good or Excellent produces noticeably cleaner transcripts than Noisy grades, particularly for proper nouns, technical terminology, and speech with moderate accents. The Clipping Detector and Frequency Response display are secondary checks that confirm gain staging is appropriate before recording begins.

What to look for

  • 25 MB per request
  • 96 dB

Opens the Microphone Quality, Noise & Latency Tester with this section's reference values shown at the top of the tool.

Open in the tool →

Noise across the frequency bins Whisper reads

Whisper's architecture pre-processes audio using log-Mel spectrogram features across 80 frequency bins from 0 to 8 kHz2. Noise that concentrates in specific bins, like HVAC rumble below 200 Hz and fan noise in the 100–400 Hz range, competes directly with the fundamental frequency range of human voice. Furthermore, Whisper processes audio in 30-second chunks. Background noise that fluctuates within those chunks can affect word boundary detection at the chunk boundaries, producing occasionally garbled words at those points.

Consistent noise across recording chunks

A Noise Floor Grade of Good or Excellent indicates consistent noise conditions throughout the recording session. Whisper processes audio in fixed 30-second segments, and the model performs best when the noise floor remains stable across those segment boundaries. If the noise level shifts dramatically mid-recording, for example when an HVAC system cycles on between segments, the model may produce different transcription quality at the boundary point compared to the surrounding audio. Running the Noise Floor Grade at the start, middle, and end of a long recording session confirms whether your room conditions remain stable throughout.

Plan long recordings around the stable reading windows the test reveals, because a chunk that straddles an HVAC cycle or traffic burst is exactly where Whisper tends to garble a word at the segment edge. Catching the instability with three readings before you press record is far cheaper than re-transcribing a compromised segment afterward, and it tells you which part of the day to avoid for critical recordings.

When latency matters for Whisper

Latency is not a factor in Whisper API batch transcription: you record first, then upload. However, if you use Whisper in a real-time pipeline (local deployment or third-party wrappers that process audio in near real-time), the Clap Latency Test becomes relevant. At high latency, real-time wrappers fall behind the audio stream and accumulate delay.

Batch use versus real-time wrappers

For pure API batch use, skip the Clap Latency Test and focus on the Noise Floor Grade and Clipping Detector before each recording session. Real-time wrappers that stream audio to a local Whisper model still depend on noise floor quality, but they add latency as a secondary concern because the model must keep pace with the incoming audio stream. If you deploy a local Whisper instance for live transcription, run both the Noise Floor Grade and the Clap Latency Test: the noise floor predicts accuracy while the latency prediction tells you how quickly the transcription will appear after each spoken phrase.

Choosing between API batch and local deployment for time-sensitive work

When transcription speed matters more than convenience, a local Whisper deployment eliminates the upload step entirely and processes audio on your own hardware, but this shifts the performance bottleneck from network latency to your CPU or GPU's inference speed. A local large-v3 model on a modern GPU transcribes in near real-time with acceptable latency, while the API introduces upload and queue delays that add seconds to minutes depending on file length and server load. CapyToolkit's pre-recording checks apply equally to both deployment paths: clean audio at the source matters more than which Whisper variant processes it.

Why clipping hurts Whisper more than noise

Clipping is more damaging to Whisper transcription than noise, because distorted waveforms create false spectral energy across multiple Mel filter banks simultaneously3. This scatters recognition confidence across multiple phoneme candidates at once, producing substitution errors rather than insertion or deletion errors: the transcribed word sounds wrong rather than being absent. The Clipping Detector confirms that your recording gain is below the saturation threshold before you begin. The Frequency Response display is useful for confirming that the microphone covers the 80 Hz–8 kHz range that Whisper's spectrogram processing uses most heavily, because a microphone that rolls off well below 8 kHz will capture the full speech spectrum while one that drops off at 3.4 kHz (the upper limit of narrowband telephone audio4) will lose critical consonant information and produce noticeably worse transcription accuracy even if the noise floor reading appears acceptable.

Setting up your recording environment for Whisper accuracy

Setting up your recording environment before a Whisper session is a one-time calibration that pays off across every recording you do in that space. Start with the Noise Floor Grade: run it three times and note the average. A consistent result within 2–3 dBFS across runs means the room is stable. Variation of 5 dBFS or more indicates an intermittent noise source (HVAC cycling, a refrigerator compressor, or traffic bursts) that will appear as noise spikes in the recording and affect word boundary detection at the 30-second chunk boundaries Whisper processes.

Assessing run-to-run variation before committing to a session

Run-to-run variation above 5 dBFS means the intermittent noise source is significant enough to create gaps in a long recording. Track which minutes of the day show stable readings and which show variation. HVAC systems typically cycle every 10–15 minutes; if your stable readings cluster in a consistent window, schedule recordings during that window rather than treating the noise floor as fixed. Three stable consecutive readings within 2 dBFS of each other confirm the room is ready.

Testing gain ceiling before a long recording session

Before any recording session intended for Whisper, run the Clipping Detector for 20 seconds while speaking at your loudest normal voice plus deliberate consonant sounds. The goal is to confirm the badge never appears during that window. If it does, reduce gain by 3 dB and repeat. Clipping discovered mid-recording means the affected segments require re-recording or produce transcript errors that are time-consuming to correct. Confirming the ceiling before starting takes under two minutes and guarantees a clean recording throughout the session.

File format and sample rate for Whisper transcription

Before recording for Whisper API upload, confirm your microphone's sample rate output in OS audio settings. Whisper internally resamples all audio to 16 kHz for its mel-spectrogram feature extraction, which means recording at 48 kHz provides no accuracy benefit beyond what 16 kHz captures. The Frequency Response display confirms whether your microphone covers the 80 Hz–8 kHz range meaningfully: a display that shows consistent energy through 8 kHz with a drop above means the signal is complete at the range Whisper uses most heavily.

File format affects the 25 MB API upload limit more than audio quality5. A 30-minute interview recorded as 48 kHz 16-bit WAV is approximately 165 MB, well above the limit. Recording as 16 kHz 16-bit WAV reduces the size to around 55 MB, still over the limit for long sessions. Recording in a compressed format such as opus or mp3 at 64 kbps reduces 30 minutes to under 15 MB while preserving accuracy, since Whisper's training included compressed audio formats at these bitrates. Before the session, pick the right Whisper audio format to prevent upload failures that require re-recording or manual file splitting.

When to use this

Run this check before any extended recording session you plan to transcribe with Whisper. Interviews, lectures, podcasts, and meeting recordings all benefit from a verified noise floor before you start. A clean input means less time correcting output.

Examples

Interview recording in an office with AC noise

Before
Noise floor −42 dBFS (Noisy) — proper nouns and technical terms misrecognized, 15% error rate
After
Moved to quieter office: noise floor −58 dBFS — error rate drops below 3%

Podcast recording with condenser mic at too-high gain

Before
Clipping badge appears during louder speech — Whisper produces artifact errors at high-energy segments
After
Input gain reduced 10 dB, no clipping badge — Whisper accuracy improves across the full episode
Sources
  1. 1.

    OpenAI, "Introducing Whisper," openai.com, September 2022. https://openai.com/research/whisper

  2. 2.

    Radford et al., "Robust Speech Recognition via Large-Scale Weak Supervision," ICML 2023, pp. 28492–28518. https://proceedings.mlr.press/v202/radford23a.html

  3. 3.

    "Clipping (audio)," Wikipedia, accessed June 2026. https://en.wikipedia.org/wiki/Clipping_(audio)

  4. 4.

    "Wideband audio," Wikipedia, accessed October 2026. https://en.wikipedia.org/wiki/Wideband_audio

  5. 5.

    Mozilla Developer Network, "MediaDevices: getUserMedia() method," developer.mozilla.org, accessed September 2026. https://developer.mozilla.org/en-US/docs/Web/API/MediaDevices/getUserMedia

FAQ

Whisper accepts mp3, mp4, mpeg, mpga, m4a, wav, and webm. The browser's MediaRecorder API defaults to webm/opus. File size is limited to 25 MB per request. For long recordings, chunk the audio before upload.

Yes, but accuracy decreases. CapyToolkit's Noise Floor Grade tells you when the recording is too noisy before you upload it. Whisper handles noise better than most models, but Noisy grade recordings are usable for casual transcription and will need more editing for technical content or proper names.

Not significantly. Whisper processes audio features at a scale where the 96 dB dynamic range of 16-bit audio is sufficient. Noise floor matters far more than bit depth for real-world transcription accuracy.

For sensitive audio (medical, legal, confidential), a local Whisper deployment does not send audio to OpenAI's servers. The API transmits audio. Both versions process audio the same way acoustically. The choice is about data privacy, not quality.

Yes. Whisper performs best on English, but supports 100+ languages. Non-English languages show higher sensitivity to noise floor levels because the training data distribution is less balanced. Excellent grades matter more for non-English transcription.

Amazon Alexa Mic Test

On single-microphone setups, Amazon Alexa's far-field design needs extra care: the Amazon Echo uses a seven-microphone circular array with beamforming, noise reduction and echo cancellation to pick up wake words across a room1. When you use Alexa through a web browser or application with a single cardioid microphone, the expected far-field behavior translates into near-field signal processing that behaves differently. Specifically, noise floor thresholds calibrated for beamforming arrays are applied to single-microphone input, which means near-mic use requires a lower noise floor than far-field expectations suggest.

Wake-word detection is the most noise-sensitive stage in Alexa's pipeline2. The keyword model that listens for "Alexa" is trained to detect a specific acoustic pattern against a background noise model. When your measured noise floor approaches −40 dBFS, false wake-word triggers increase and genuine wake words become less reliable. Furthermore, the Clap Latency Test reveals the latency that affects how quickly Alexa's command processing feels after the wake word is detected.

What to look for

  • below -50 dBFS
  • below 60 ms
  • above 100 ms

Opens the Microphone Quality, Noise & Latency Tester with this section's reference values shown at the top of the tool.

Open in the tool →

Wake-word false triggers and room noise

Alexa's wake-word detection uses an always-listening keyword model that runs independently of the main speech recognition pipeline. This model has its own noise floor sensitivity curve: false activations increase significantly when ambient noise reaches the −40 dBFS range (Noisy on the Noise Floor Grade)3. Furthermore, genuine wake-word activations become unreliable above −45 dBFS because the detection confidence threshold requires a certain speech-to-noise ratio to trigger.

Practical minimum for single-mic wake words

A Noise Floor Grade of Good (below −50 dBFS) is the practical minimum for reliable single-microphone wake-word detection. Echo devices compensate for noisier environments using microphone array beamforming that a single cardioid cannot replicate. Without that array processing, the single microphone needs a cleaner signal to achieve the same wake-word confidence. If your Noise Floor Grade is in the Noisy range, the wake-word model may fail to detect your voice from more than one meter away, or it may trigger falsely on background sounds that briefly resemble the Alexa phoneme pattern.

Command response delay and speaker coupling

After wake-word detection, Alexa transitions to command processing mode. Round-trip latency measured by the Clap Latency Test contributes to the perceived delay between the end of your command and Alexa's response. Unlike conversational AI services, Alexa's responses are typically brief and discrete, so latency above 100ms does not degrade the experience as severely. Nevertheless, keeping the Clap Latency Test result in the Good range ensures that your command reaches Amazon's servers with minimal browser-side delay. This matters more for multi-step interactions where you issue several commands in quick succession and need each one to register before the next begins processing.

Echo Loopback for wake-word coupling

Building on this, the Echo Loopback test is particularly useful for Alexa: playing Alexa's audio response through speakers near the microphone creates acoustic feedback that can trigger false wake words. The Loopback test reveals how much of the speaker output your microphone is capturing. If the recorded level during Alexa playback exceeds −45 dBFS relative to your voice level, the acoustic coupling between speaker and microphone is strong enough to cause false activations during normal use. Repositioning the microphone to maximize the cardioid's rear rejection toward the speaker, or switching to headphones for Alexa monitoring, eliminates this coupling entirely.

Clipping in the band the Alexa wake word uses

Alexa's wake-word model is optimized for the 100 Hz–4 kHz range where the "Alexa" phoneme pattern lives4. Clipping during command words creates distortion that spreads across this band, degrading phoneme matching. The Frequency Response display helps confirm that your microphone covers the 100 Hz–4 kHz band consistently; a significant dip in the 1–3 kHz region will affect wake-word detection reliability. For browser-based Alexa use, the command is typically one to three words, making clipping on any single command word more consequential than in longer conversational speech.

Because the wake-word model evaluates the entire command in one brief analysis window, a clipped phoneme at the start of the command can corrupt the model's confidence score for the entire phrase, causing Alexa to either misinterpret the command or fail to respond at all. Running the Clipping Detector while speaking a few test commands at your normal volume before relying on Alexa for important tasks ensures that the gain staging you have set preserves the full phoneme pattern the wake-word model needs.

Near-field microphone positioning for single-mic Alexa use

Near-field positioning for single-mic Alexa use differs substantially from the far-field assumptions the Echo device's array is designed for. An Echo placed on a shelf captures your voice from 1–3 meters using beamforming that rejects ambient noise. A single cardioid microphone on a desk functions best at 15–30 cm, where the direct voice signal dominates over room noise without pushing the Clipping Detector during normal command volume. Moving further away forces higher gain, which raises the measured noise floor and degrades wake-word reliability at the same time.

The cardioid pattern rejects sound from the rear null point by 15–20 dB relative to on-axis sensitivity5. Position the microphone so that the rear of the capsule faces the Echo speaker, rather than placing the capsule between you and the speaker. This reduces the speaker's playback from entering the Noise Floor Grade measurement and lowers the risk of Alexa's own voice triggering false wake-word detection. The Frequency Response display confirms whether speaker audio is still coupling into the microphone after repositioning: if the display shows energy during Alexa's response that was absent during silence, coupling is still present and further adjustment is needed.

Testing Alexa playback as a noise source with the Echo Loopback

Alexa's synthesized voice output, played through a nearby speaker, contributes to your microphone's noise floor measurement and can trigger false wake-word detections during playback. The Echo Loopback test captures this accurately: run the test while Alexa is speaking its response to a command. The playback recording reveals exactly how much of Alexa's voice the microphone captures. If the recorded level during Alexa's playback exceeds −50 dBFS relative to the signal level during your commands, the coupling between speaker and microphone is strong enough to trigger false activations.

Repositioning versus headphones to eliminate acoustic coupling

Reducing speaker volume lowers the coupling level proportionally but may reduce playback intelligibility. Repositioning the microphone to maximize the cardioid's rejection null toward the speaker is more effective. Switching to headphones for Alexa monitoring eliminates the speaker-to-microphone coupling entirely and produces the cleanest Noise Floor Grade result. The Echo Loopback test is the most direct way to quantify how much coupling exists before and after a repositioning change: run the test during an Alexa response, note the captured level, reposition, and run the test again to confirm the improvement numerically.

To quantify Alexa voice coupling in the loopback before relying on wake words, use headphones when repositioning is not practical, because they remove the speaker entirely from the microphone's acoustic environment rather than merely reducing its level. If you cannot use headphones, aim for the −50 dBFS captured-level threshold as the cutoff where false triggers become likely, and treat any reading above it as a signal to move the microphone further from the Echo speaker or angle the rear null toward it.

When to use this

Use this check when setting up Alexa web or app access for the first time, or when Alexa frequently fails to wake or misunderstands commands in your current environment.

Examples

Laptop microphone in a shared office

Before
Noise floor −38 dBFS — Alexa triggers on colleague conversations, misses some genuine wake words
After
External USB cardioid microphone: −60 dBFS — reliable wake detection, no false triggers

Desktop microphone with speakers nearby

Before
Echo Loopback shows Alexa's voice response bleeding back into the mic — false wake-word triggers during playback
After
Moved mic away from speaker, angled cardioid to reject speaker: echo feedback eliminated
Sources
  1. 1.

    Amazon, "Amazon Makes the High-Performance 7-Mic Voice Processing Technology from Amazon Echo Available to Third-Party Device Makers," press.aboutamazon.com, April 2017. https://press.aboutamazon.com/2017/4/amazon-makes-the-high-performance-7-mic-voice-processing-technology-from-amazon-echo-available-to-third-party-device-makers

  2. 2.

    Yixin Gao et al., "On Front-end Gain Invariant Modeling for Wake Word Spotting," arxiv.org, 2020. https://arxiv.org/abs/2010.06676

  3. 3.

    Springer, "Speech Recognition in Adverse Conditions," link.springer.com, 2026. https://link.springer.com/article/10.1186/s13636-026-00458-1

  4. 4.

    Wikipedia, "Formant," en.wikipedia.org, accessed June 2026. https://en.wikipedia.org/wiki/Formant

  5. 5.

    Shure, "515SA/B Guide," pubs.shure.com, accessed June 2026. https://pubs.shure.com/view/guide/515SA-B/en-US.pdf

FAQ

Echo devices use a microphone array with beamforming, and CapyToolkit's Noise Floor Grade shows why a single cardioid microphone needs more margin. A single cardioid microphone cannot reject diffuse room noise the same way. Keep the grade below −50 dBFS to compensate for the absent array processing.

Yes. Alexa web access uses getUserMedia, which works with any browser-accessible microphone. Reliability improves significantly with a dedicated cardioid microphone versus a built-in laptop microphone.

Alexa's synthesized voice contains phoneme patterns that can briefly resemble the wake word. The Echo Loopback test shows how much of the playback your microphone captures. If the level is above −50 dBFS of the original signal, false triggers are more likely.

Below 60ms produces the fastest command response. Between 60 and 100ms is acceptable. Above 100ms is Noticeable, which makes Alexa feel sluggish compared to a dedicated Echo device.

Yes, but Bluetooth adds 40 to 120ms to browser-side latency. For frequent command use, a wired microphone produces noticeably faster responses.

Apple Siri Mic Test

With Apple Silicon, Siri and macOS Dictation both use on-device speech recognition models, processing audio locally without a server round-trip for supported queries1. On-device processing has a distinct noise floor characteristic: the recognition model running at inference time on the Neural Engine processes audio in real time without the error-correction passes possible in batch transcription systems2. This makes the on-device model sensitive to clipping in a specific way: a single clipped frame at the wrong moment corrupts a phoneme that the real-time decoder cannot recover from by looking at subsequent context.

Apple's voice-processing audio unit includes automatic gain control for the processed microphone signal3. Gain control normalizes amplitude variation rather than removing background noise, so it does not lower the noise floor. Consequently, the Noise Floor Grade reflects the true hardware noise floor even with OS gain control active, while the Clipping Detector may show cleaner results than the hardware gain setting implies, because gain control is pulling the level down before peaks reach 0 dBFS.

What to look for

  • Good grade, below -50 dBFS
  • Excellent grade, below -60 dBFS
  • 16 kHz, focused on the 80 Hz to 8 kHz range
  • about 5 to 10 dBFS stricter for continuous Dictation

macOS gain control evens out level but does not remove background noise, so a clean Clipping Detector reading does not by itself confirm correct gain staging.

Opens the Microphone Quality, Noise & Latency Tester with this section's reference values shown at the top of the tool.

Open in the tool →

Hey Siri on a single desk microphone

Siri and macOS Dictation use always-on listening for "Hey Siri" when enabled, with the same far-field detection limitations as Alexa in browser contexts: a single cardioid microphone requires a lower noise floor than Apple assumes for the microphone array in an iPhone. A Noise Floor Grade of Good or Excellent (below −50 dBFS) is the threshold for reliable detection. Furthermore, macOS Dictation's continuous mode (where speech is transcribed as you speak) shows more sensitivity to noise floor than command-based Siri, because each word must be resolved without a natural pause that aids segmentation.

Run CapyToolkit's Noise Floor Grade before extended dictation, because the continuous mode needs a cleaner floor than short Siri commands. In command mode, Siri processes a brief 2 to 5 word phrase with a constrained vocabulary that the acoustic model can resolve even with moderate background noise. Continuous Dictation, by contrast, transcribes unlimited vocabulary in a real-time stream where every word must be resolved without the benefit of command-boundary context. The practical difference is approximately 5 to 10 dBFS: a Noise Floor Grade of Good may suffice for Siri commands, but continuous Dictation benefits from Excellent grades to maintain accuracy across long passages.

On-device processing and browser latency on macOS

Siri on macOS performs voice processing locally for many commands, which eliminates the server-side network latency that affects cloud-based AI services4. The Clap Latency Test therefore measures only browser-to-OS audio path latency for browser-based access scenarios. For dictation in native macOS applications, latency depends on Apple's Core Audio stack rather than the WebAudio path measured here.

Browser latency versus native dictation latency

If the Clap Latency Test shows high values in Chrome on macOS, Safari typically shows lower latency on Apple hardware because it uses Core Audio APIs more directly. The browser audio sandbox that Chrome applies adds processing stages that increase round-trip delay compared to Safari's more direct path to the audio hardware. For the lowest latency browser-based audio testing on macOS, Safari is the better choice. However, native macOS dictation bypasses the browser audio stack entirely and uses Core Audio directly, which means the latency you experience during actual Dictation use will be lower than what the Clap Latency Test in any browser reports.

Clipping and the 16 kHz speech model

Apple's Neural Engine based speech recognition processes audio at 16 kHz sample rate, focusing model capacity on the 80 Hz to 8 kHz range5. OS gain control partially masks gain overload by lowering the level before it reaches the digital ceiling, but if input gain is so high that even the reduced peak exceeds 0 dBFS, the Clipping Detector will flag it. Furthermore, the Frequency Response display is useful for confirming that your microphone's response through 80 Hz to 8 kHz is consistent, because the on-device model is not designed to compensate for severe frequency response gaps in the input.

Why OS gain control does not replace proper gain staging

macOS gain control operates on the digital audio stream after the ADC has already captured and quantized the signal, which means it can reduce peak amplitudes but cannot recover headroom that was lost during the analog-to-digital conversion stage. If your input gain drives the ADC into saturation before gain control even receives the samples, it simply reduces the amplitude of already-clipped digital values without restoring the original waveform shape. Setting the input gain low enough that gain control rarely has to step in preserves the full dynamic range that the on-device model needs for accurate phoneme discrimination.

Run the Clipping Detector before relying on Dictation for important text, because a clean reading after gain control is not proof the gain is correct: it only means the processed output stayed under the ceiling, while the underlying signal may already be saturated and distorted at the conversion stage. A clean detector reading combined with a Noise Floor Grade in the Excellent range confirms the input chain is healthy end to end rather than merely flattened by the OS processor.

macOS Dictation versus Siri: different noise floor requirements

Between macOS Dictation's continuous transcription mode and Siri's discrete command processing, the noise floor requirement differs by approximately 5 to 10 dBFS. Siri processes short commands, typically 2 to 5 words, and has a brief, defined recognition window for each interaction. The acoustic model can use surrounding context to resolve a word partially obscured by noise because the command vocabulary is constrained. A Noise Floor Grade of Good (below −50 dBFS) is sufficient for reliable Siri command recognition with common vocabulary in a moderately quiet room.

Why continuous Dictation requires a higher grade than Siri commands

Continuous Dictation mode transcribes speech in a stream without natural command boundaries. The acoustic model encounters longer, less predictable vocabulary including proper nouns, technical terms, and complex sentence structures. Every word must be resolved on its own merits without the contextual disambiguation that short command recognition relies on. For Dictation accuracy that approaches what you would expect from an accurate typist, a Noise Floor Grade in the Excellent range (below −60 dBFS) is the practical target. Background noise in the Noisy range produces Dictation accuracy that requires frequent manual correction, negating the time benefit Dictation is supposed to provide.

Setting the correct input device in macOS Audio MIDI Setup

macOS routes microphone audio through the System Settings > Sound > Input panel, but the Audio MIDI Setup application provides more granular control over which device is selected and at what sample rate. When multiple microphones are connected, the default input in System Settings controls what most applications use. Siri and Dictation respect the default input device unless a specific application overrides it. Verifying that the Noise Floor Grade test and Siri are using the same microphone requires checking System Settings > Sound > Input and confirming your preferred device is selected before testing.

If the Noise Floor Grade returns an Excellent result but Siri accuracy remains poor, the two tools may be using different devices. Some third-party applications set themselves as the default input and do not release it properly after closing. Check System Settings > Sound > Input after closing all other applications to confirm the correct device is active. Running the Noise Floor Grade immediately before asking Siri a question lets you see which mic Siri actually uses and confirms both are operating on the same audio path within the same OS session, making the pre-session check a reliable calibration step rather than a theoretical one.

When to use this

Run this check before using Siri voice activation or macOS Dictation with an external microphone, especially when switching between the built-in microphone and an external USB or Bluetooth device on macOS.

Examples

External USB condenser on macOS with boosted OS gain

Before
OS gain control masking clipping — Dictation still produces errors on plosive sounds (p, b, t, k)
After
Reduced OS gain by 15 dB, gain control no longer working hard: accurate Dictation across all consonants

Wireless AirPods used as microphone on macOS

Before
Clap Latency Test: 110ms (High) — Siri command recognition feels delayed
After
Switched to USB wired microphone: 35ms — Siri responds instantly
Sources
  1. 1.

    Apple, "Using On-Device Speech Recognition," developer.apple.com, 2019. https://developer.apple.com/videos/play/wwdc2019/256/

  2. 2.

    Apple, "Voice Trigger System for Siri," machinelearning.apple.com, accessed June 2026. https://machinelearning.apple.com/research/voice-trigger

  3. 3.

    Apple, "kAUVoiceIOProperty_VoiceProcessingEnableAGC," developer.apple.com, accessed October 2026. https://developer.apple.com/documentation/audiotoolbox/kauvoiceioproperty_voiceprocessingenableagc

  4. 4.

    The Verge, "Apple Siri On-Device Speech Recognition," theverge.com, 2021. https://www.theverge.com/2021/6/7/22522993/apple-siri-on-device-speech-recognition-no-internet-wwdc

  5. 5.

    Apple, "Hey Siri: An On-device DNN-powered Voice Trigger," machinelearning.apple.com, accessed June 2026. https://machinelearning.apple.com/research/hey-siri

FAQ

Gain control lowers the level before the digital ceiling, but if input gain is high enough that peaks still clip, the Clipping Detector will fire. Set gain low enough that gain control is not constantly working.

No. macOS Siri uses a different model variant optimized for keyboard-and-voice interaction patterns. The audio requirements are similar, but the language model component differs between platforms.

Safari uses Apple's Core Audio APIs more directly and bypasses some of the browser audio sandbox overhead that Chrome applies. For lowest-latency audio on macOS, Safari is typically faster in WebAudio benchmarks.

Significantly, in typical office environments. The MacBook built-in microphone is designed for omnidirectional pickup from any direction. CapyToolkit's Noise Floor Grade helps confirm that an external cardioid aimed at your mouth produces a better grade and more accurate recognition.

On Apple Silicon Macs with macOS Ventura and later, Dictation processes audio on-device using the Neural Engine. Some complex Siri requests may still use server-side processing. For Dictation specifically, on-device processing is the default.

FAQ

The Noise Floor Grade. Live assistants decide when you start and stop talking by comparing your voice with the room, so a quiet room fixes more problems than any setting. Add the Clipping Detector for anything that transcribes your speech.

Its turn detector heard enough silence, or treated a pause as the end of your turn. Background noise makes this worse in both directions. Get the grade to Good or better and avoid long pauses mid-sentence.

It works, but it adds delay and, while the headset microphone is active, often switches to a narrowband mode that sounds like a phone call. A wired or USB microphone gives the assistant a cleaner, faster signal.

For assistants you can interrupt, yes if your speakers are near the microphone. The assistant's own reply can reach your microphone and count as you speaking. CapyToolkit's Echo Loopback lets you hear how much playback the microphone picks up.

No. CapyToolkit doesn't upload anything from these tests; your browser measures the microphone locally. The assistants themselves differ: Claude Code dictation streams audio to Anthropic for transcription, and Mac Dictation can run on your device depending on the setting.

The highest setting where your loudest normal speech never triggers the Clipping Detector, then a few steps lower for headroom. Recheck it whenever you change rooms or plug in a different microphone.

Additional resources