PuppyIP Resource Center
AI Tool Updates 10 minutes Published 2026-08-28

Choosing Gemini 3.5 Transcribe: Live's 10-minute limit, speakers and timestamps

If you plan to use Gemini 3.5 Transcribe for live captions, meeting notes or support recordings, do not choose a model simply by whether it is real-time. Google released file and Live endpoints on August 26. Live sessions last at most 10 minutes and lack speaker diarization and word-level timestamps. Enabling either on the file endpoint reduces its limit from one hour to 30 minutes. First decide whether real-time output, speaker identification and word-level positioning are required.

Gemini 3.5 Transcribe Gemini API Live captions Speaker diarization Speech transcription

Service eligibility and regional restrictions

PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.

Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.

Key Takeaways

  • Google released gemini-3.5-transcribe and gemini-3.5-transcribe-live on August 26, 2026. The official model page lists them as stable speech-to-text models.
  • The Live endpoint provides low-latency streaming transcription through Live API, with a maximum 10-minute session and no speaker diarization or word-level timestamps.
  • The file endpoint accepts up to one hour of audio per request, or 30 minutes with speaker diarization or word-level timestamps enabled.
  • Smart transcription cleans up speech fillers and formatting, but is incompatible with timestamps and speaker mode. Prefer Verbatim when auditing exact speech.
  • Validate language, terminology, speaker roles, latency and cost with representative audio. Stop if required features are absent, timestamps drift or segmentation loses context.

Live captions work, but meeting review lacks speakers and a timeline

A developer sees “Live” and sends a meeting or stream longer than ten minutes over WebSocket. Captions appear for the first few minutes, then the session reaches the 10-minute boundary. Exported text has neither speaker labels nor word-level timestamps. Symptoms include reconnects, caption gaps, mixed speakers and an inability to locate speech in the recording. Costs include missing live words, manual rearrangement, unverifiable support quality reviews and incorrect attribution of meeting responsibilities.

The wrong response is to keep extending the Live session, prompt for unsupported fields or split every recording into ten-minute pieces. First list three hard requirements: must it be real-time, distinguish speakers or locate every word? If either of the last two is required, evaluate the file endpoint first.

The endpoints solve different tasks

gemini-3.5-transcribe-live handles low-latency input such as microphones, broadcasts and voice agents, transcribing through Live API. The official limit is 10 minutes per session, without word-level timestamps or speaker diarization. It suits immediately readable captions, but cannot provide a full speaker and word-level audit trail.

gemini-3.5-transcribe handles prerecorded files through Files API and Interactions API. It supports speaker diarization, word-level timestamps, up to 1,000 custom vocabulary entries and Smart transcription. Ordinary audio can last up to one hour; enabling either diarization or word-level timestamps lowers the limit to 30 minutes.

Choose by the deliverable, not the model name

For live captions, in-call prompts or voice-agent input, start with Live and implement controlled continuation before ten minutes. For meeting notes, interviews, support quality reviews and legal or compliance checks requiring speakers and timelines, prefer the file endpoint, keep segments within 30 minutes and retain original audio. Offline verbatim transcripts without speaker labels or word timestamps can use the one-hour file limit.

Google says the model automatically detects more than 85 language locales and supports language switching within a session. That does not establish acceptable quality for every Chinese accent, industry term or noise condition. Public documentation gives no Chinese accuracy rate or regional quality ranking. Test representative samples yourself; do not invent WER figures or experience claims.

Choosing Smart, Verbatim, timestamps and speakers

For directly readable captions or notes, test Smart transcription, which handles fillers, repetition and formatting. To preserve exact speech, compare word by word, or use timestamps and speaker labels, choose Verbatim. Official documentation says Smart mode is incompatible with timestamp_granularities and diarization_mode; do not force them into one request.

The file endpoint distinguishes up to eight speakers, but attribution for three or more remains experimental. You can supply up to 1,000 vocabulary phrases, though Google says fewer than 100 typically work best. Start with product names, people and abbreviations likely to be misrecognized rather than inserting an entire dictionary.

Seven acceptance checks before release

First prepare quiet, noisy, accented, multiple-speaker and code-switching samples from real tasks. Second choose Live or file. Third segment within feature limits and retain original time offsets. Fourth add only necessary terminology. Fifth compare against a human reference for missing words, numbers, proper names, speakers and timing. Sixth record first-word latency, total duration, input/output tokens and failed retries. Seventh test segment merging, ten-minute continuation and recovery.

An official description or one demo is not acceptance testing. Simulate reconnects and network jitter for Live; test the 30-minute boundary, overlapping speakers and silence for files. For audit or public output, retain human review and links back to original audio.

Pricing, data use and capability limits

Google's pricing page lists paid file transcription at $2 per million audio-input tokens and $12 per million text-output tokens, estimated at about $0.005 per minute. Live costs $3.50 input and $21 output, estimated at about $0.009 per minute. A free tier is available, but its data may be used to improve products; the paid tier is marked as not used for product improvement. Prices and eligibility may change, so recheck project billing before release.

Transcribe is not general audio understanding, speech synthesis or real-time translation. The file model does not support caching, function calling, file search, thinking, Batch, Flex or Priority; Live does not support Google Search grounding. Sharing the Gemini name does not imply every other model's capabilities.

Stop, rollback and troubleshooting conditions

Stop expansion and return to the last stable transcription path if sessions repeatedly develop gaps, speaker attribution is unacceptable, important numbers or names are often wrong, timestamps drift, merged segments lose context, or cost or latency exceeds limits. Preserve model ID, mode, language hints, audio duration, feature settings, request ID and redacted failing clips before reducing variables one by one.

For 401, 403, 429 and 503, check keys, regional access, quotas and capacity respectively using the Gemini API 429 and 503 guide. Use the proxy connection checklist only with clear DNS, TLS, connection-timeout or proxy-authentication evidence. For stable access to API documentation and consoles, visit the PuppyIP website to learn about network environments. A network cannot change model capabilities, quotas or regional eligibility.

Common mistakes: real-time does not mean more features

The first mistake is assuming Live outputs speaker labels and word-level timestamps; it explicitly does not. The second is sending a 60-minute recording with diarization after seeing the one-hour file limit; that configuration allows 30 minutes. The third is treating Smart output as verbatim evidence when it cleans up fillers and formatting. The fourth is claiming equal quality across all 85-plus language locales.

The fifth is blaming every disconnect on a proxy. Live has an explicit 10-minute session limit and files have feature-dependent duration limits. Check product boundaries and request configuration before investigating network evidence, avoiding pointless retries.

Sources

Frequently Asked Questions

How long can a Gemini 3.5 Transcribe Live session run?

Google's model page specifies at most 10 minutes per Live session. Long-running captions need controlled continuation and checks for missing or repeated words at reconnects.

Does Live support speaker diarization and word-level timestamps?

No. These are available only on the file endpoint. Evaluate gemini-3.5-transcribe when speakers and word-level timelines are required.

Can the file endpoint process one hour of audio?

Ordinary file transcription allows one hour. Enabling speaker diarization or word-level timestamps lowers the limit to 30 minutes.

Can Smart transcription be combined with timestamps?

No. Smart mode is officially incompatible with timestamp_granularities and diarization_mode. Use Verbatim when auditing exact speech.

How many speakers can it distinguish?

The file endpoint supports up to eight speakers, but attribution for three or more remains experimental and needs real multiple-speaker acceptance samples.

Is a disconnect after ten minutes a proxy problem?

Not necessarily. Check the Live 10-minute session limit first, then error codes and connection evidence. A proxy cannot change official duration or feature limits.