Free speech APIs — transcription and text-to-speech

Free speech-to-text and text-to-speech APIs, from free monthly minutes to hundreds in credit.

Audio APIs — transcription (STT) and synthesis (TTS) — often ship a real free tier or a sizeable one-time credit. The providers below offer free audio/speech access, from a few free hours a month to hundreds of dollars in starting balance.

Check the "catch" on each provider's page: some free speech tiers forbid commercial use or watermark the output.

Top pick — Groq

Lowest-latency inference for open-weight models Details →

26 verified providers

ProviderFree tierRate limitsGotchas
Google Gemini API (AI Studio)Gemini 2.5 Flash, 2.5 Flash-Lite, 2.5 Pro (limited), embeddings, TTS modelsPer-model; Google no longer publishes the free-tier numbers in its public docs — they are shown only on the AI Studio rate-limit page after sign-inno cardno phonecommercial OKOpenAI-compatible
GroqOpen-weight models (GPT-OSS, Qwen) plus Whisper, no credit card requiredFree plan examples: openai/gpt-oss-120b, openai/gpt-oss-20b and qwen/qwen3.8-27b: 30 RPM/1K RPD/8K TPM/200K TPD; Whisper models: 20 RPM/2K RPD/7.2K audio seconds/hour and 28.8K/day; limits vary by modelno cardphone requiredcommercial OKOpenAI-compatible
Cloudflare Workers AI10,000 Neurons/day, all account plans10,000 Neurons/day (free allocation). Per-task request limits: text generation 300 requests/min (models that require Workers Paid: 20 requests/min per model); text embeddings 3,000/min; speech recognition, text-to-image, translation, image-to-text 720/minno cardno phonecommercial OKOpenAI-compatible
HuggingFaceFree CPU Basic + ZeroGPU for Spaces; Inference Providers has no included credit on the Free plan (credits must be purchased) — $2.00/mo is included only on PRO/Team/EnterpriseNo RPM/TPM published, only credit amountsno cardOpenAI-compatible
OVHcloud AI EndpointsSeven models are currently listed as Free: two Qwen3Guard moderation models, Stable Diffusion XL and four NVIDIA Riva TTS voices; access is available anonymously or with an API key tied to a Public Cloud projectAnonymous access: 2 requests/min per IP per model. Authenticated (API key): 400 requests/min per project per model. Exceeding either returns HTTP 429no cardOpenAI-compatible
Scaleway Generative APIs1,000,000 tokens free + 60 min Whisper transcription; billing starts at token 1,000,001With a validated payment method: 300 RPM / 200k TPM per model; 600 RPM and higher TPM once identity is verifiedno cardOpenAI-compatible
Deepgram$200 free credit on signup (no card, no expiration) — Nova speech-to-text and Aura text-to-speech at pay-as-you-go ratesSTT pre-recorded up to 50 concurrent; STT streaming up to 150; TTS REST up to 15; TTS streaming up to 45no cardno phonecommercial OK
AssemblyAI$50 free credit on signup (no card) — pre-recorded & streaming speech-to-text, Speech Understanding, and an OpenAI-compatible LLM Gateway (25+ models)Free tier: 5 parallel transcriptions; streaming 5 new streams/min; global 20k requests / 5 minno cardno phonecommercial OKOpenAI-compatible
Pollinations.aiQuest Pollen: create an account, complete eligible Quests on the Quests dashboard and claim the rewards; Quest Pollen is spent before paid Pollen on regular (non-paid-only) models. No anonymous/no-key generation — every generation endpoint requires an API key and costs Pollen ($1 ≈ 1 Pollen). Text, image, video, audio, embeddings and 3D models behind one OpenAI-compatible APINo per-request rate limits published for sk_ keys; usage is bounded by the Pollen balance (legacy raw pk_ keys: 1 Pollen per IP per hour)no cardno phoneOpenAI-compatible
Twelve Labs (Marengo Embed)Free plan: the Marengo 3.5 Embed API (video, audio, image, text and document inputs) and the Pegasus Analyze API at no cost, plus 600 minutes (10 hours) of video indexing in totalEmbed video/audio: 3,000 RPD, 25 RPM; embed text/image: 3,000 RPD, 600 RPMno card
SpeechmaticsFree plan: $100 credit grant on signup (no payment card) for Speech-to-Text; Text-to-Speech lists the first 1 million characters freeFree: 2 concurrent real-time sessions; Batch API: 10 new jobs/s and 50 job-status requests/s (bounded by the $100 credit)no card
Speechify API500K characters/month TTS (hard cap; pauses until next month), catalog voices, streaming and SSMLFree: 3 concurrent requests, 1 request/s sustained (burst 10); hard monthly cap, no top-upsno cardno phonecommercial OK
Hume AI (Octave TTS)Free tier for Octave TTS and EVI; new accounts start with $20 in creditsRate set by subscription tier, numbers not published; max 5,000 characters per utterance and 5 generations per requestno cardeval only
Unreal Speech250,000 characters/month TTS (~6 hours of audio)No rate limit published in the current V8 API docs (legacy v7 docs listed 2 requests/second on Free). Per-request size: /stream up to 1,000 characters, /speech up to 3,000, /synthesisTasks up to 500,000no cardcommercial OK
ElevenLabs10,000 credits/month shared across Text-to-Speech, Speech-to-Text and more (~10 min TTS/month)2 concurrent requests on the Free planno cardno phoneeval only
Sarvam AI₹100 in free credits on signup, usable across any API including chat completion (Sarvam-105B) and speech (STT/TTS); credits never expireStarter plan: 40 req/min for chat with Sarvam-105B (60 for other default models), STT REST and TTS REST (bulbul:v3: 30 req/min); 20 concurrent STT WebSocket streams; bounded by the ₹100 creditcommercial OKOpenAI-compatible
Gladia€50 in free credits on signup for speech-to-text (~80+ hrs pre-recorded or 60+ hrs real-time at current rates)30 real-time and 25 async concurrent requests (Starter pay-as-you-go plan that carries the €50 credit)no card
RimeFree minutes for every new account (Starter plan): the pricing card says ~800 minutes (~800k characters), while the FAQ and docs say 3,000Starter: 20 concurrent TTS generationsno card
Cartesia20,000 credits/month (~27 min of Sonic TTS or ~1h51 of Ink speech-to-text)2 concurrent TTS requests, 8 concurrent STT; 20,000 credits/monthno cardeval only
Fish AudioApp Free plan: 8,000 credits/month (~7 min, up to 500 characters per generation). API: the s2.1-pro-free TTS model is $0 through 2026-11-30 under fair-use limits (pay-as-you-go API, no subscription)API: 5 concurrent requests (Starter tier, <$100 paid); s2.1-pro-free under unpublished fair-use limits. App Free plan: 500 characters per generation, 8,000 credits/monthno cardeval only
Camb.aiFree plan: 5,000 credits one-time (no monthly renewal) across TTS, dubbing, speech and translation toolsNot published by the provider; bounded by the one-time 5,000-credit grant
Rev AIFree credits equivalent to 5 hours of Reverb ASR (speech-to-text), usable across all Rev AI productsBounded by the 5-hour credit grantno card
Voicegain$50 in free credits on signup (no credit card) — speech-to-textDefault for new accounts: 75 API requests/minute and 2,000/hour; 4 concurrent real-time ASR requests; OFF-LINE transcription up to 4 hours of audio per hour (queue of 10 jobs)no card
Smallest.ai (Waves)$10 in free credits with full access on Pay As You Go — TTS (Lightning v3.1), STT (Pulse) and voice agents; speech-to-speech (Hydra) is beta/request-accessPay As You Go: 15 concurrent TTS streams, 100 TTS requests/min, 450M chars/month cap; 100 concurrent STT streams; 20 concurrent agent calls (bounded by the $10 credit)
Retell AI$10 in free trial credits (once per email address) plus 20 concurrent calls included — voice-agent orchestration (STT + LLM + TTS)20 concurrent calls per workspace (Pay-As-You-Go default)no card
Typhoon (SCB 10X)Free to use research showcase API — all Typhoon models at $0typhoon-v2.5-30b-a3b-instruct: 5 requests/second, 200 requests/minute; typhoon-ocr: 2 requests/second, 20 requests/minute; typhoon-asr-realtime: 100 requests/minuteno cardno phonecommercial OKOpenAI-compatible

FAQ

Is there a free speech-to-text API?

Yes — Groq (Whisper), Cloudflare Workers AI and others offer free STT; several dedicated speech providers also grant a sizeable one-time credit.

Which free speech API is best for production?

Check each provider’s rate limits and whether commercial use is allowed — some free speech tiers are evaluation-only or watermark their output.

More guides

← All guides