Free OpenAI-compatible APIs (drop-in base_url)

Point the OpenAI SDK at a different base_url and keep all of your existing code.

An OpenAI-compatible API exposes the same /chat/completions shape, so switching is usually a one-line change: swap base_url and the key. That makes these the lowest-friction way to move side projects off paid OpenAI usage, or to add a free fallback behind the same SDK.

Each provider's exact base URL is on its page, and the free models you can pass as model are in the model index.

Top pick — Typhoon (SCB 10X)

Thai-language LLMs, OCR and speech via an OpenAI-compatible research API Details →

35 verified providers

ProviderFree tierRate limitsGotchas
Google Gemini API (AI Studio)Gemini 2.5 Flash, 2.5 Flash-Lite, 2.5 Pro (limited), embeddings, TTS modelsVaries by model: 5-30 RPM and 15-1,000 RPD (e.g. 2.5 Pro: 5 RPM/100 RPD; 2.5 Flash: 10/250; 2.5 Flash-Lite: 15/1,000; embeddings: 100 RPD; TTS: 15 RPD)no cardno phonecommercial OKOpenAI-compatible
GroqOpen-weight models (Llama, Qwen, GPT-OSS) plus Whisper, no credit card requirede.g. llama-3.1-8b-instant: 30 RPM/14.4K RPD/6K TPM/500K TPD; llama-3.3-70b-versatile: 30 RPM/1K RPD/12K TPM/100K TPD; qwen/qwen3.6-27b: 30 RPM/1K RPD/8K TPM/200K TPD; similar for GPT-OSS and Whisper modelsno cardphone requiredcommercial OKOpenAI-compatible
OpenRouterA rotating set of models with a :free suffix (~14 today; count fluctuates), single API across many providers20 req/min; 50 req/day under 10 credits purchased lifetime, 1000 req/day once 10+ credits purchased (one-time, not a subscription)no cardno phonecommercial OKOpenAI-compatible
Cloudflare Workers AI10,000 Neurons/day, all account plans30+ models: LLMs (Llama, Mistral, DeepSeek, Qwen...), embeddings, image, audiono cardno phonecommercial OKOpenAI-compatible
CohereTrial (evaluation) API keys covering chat, embed and rerank1,000 API calls/month total; 20 req/min chat; 2,000 inputs/min embed; 10 req/min rerankno cardno phoneeval onlyOpenAI-compatible
Cerebras$5 in free credits for new accounts, usable across all public modelsPublished Free Trial per-model limits: 5 RPM / 30,000 TPM / 1,000,000 TPH / 1,000,000 TPD (e.g. gpt-oss-120b, zai-glm-4.7, gemma-4-31b); limits vary by modelcard requiredcommercial OKOpenAI-compatible
Mistral (La Plateforme)"Restrictive" free tier explicitly for "try and explore" — official docs say to upgrade for "actual projects and production use"Not published publicly; exact caps only visible in-console after login (admin.mistral.ai)no cardphone requiredeval onlyOpenAI-compatible
HuggingFaceFree CPU Basic + ZeroGPU for Spaces; Inference Providers has a monthly credit ($0.10/mo on Free plan, $2.00/mo on PRO/Team/Enterprise)No RPM/TPM published, only credit amountsno cardOpenAI-compatible
SiliconFlowSeveral models permanently free (e.g. Qwen2.5-7B-Instruct and others) at $0 cost, plus a $1 welcome credit for paid modelsFixed per-model limits for free models; generic docs cite ranges of 1,000-10,000 RPM and 50,000-5,000,000 TPM depending on model tier — exact limits shown in-accountno cardphone requiredOpenAI-compatible
Z.ai (Zhipu AI / GLM)GLM-4.5-Flash, GLM-4.7-Flash (text), and GLM-4.6V-Flash (vision) are officially listed as $0 cost (input, cached input, and output) on a permanent basisNot specified with concrete RPM/TPM figures in public docsno cardno phonecommercial OKOpenAI-compatible
IBM watsonx.ai (Lite plan)Lite plan: 300,000 tokens/month for foundation model inference, 20 CUH/month for ML tooling, 100 pages/month of document text extraction2 inference requests per second (explicitly documented for the Lite plan)card requiredOpenAI-compatible
OVHcloud AI EndpointsTwo Qwen3Guard models (Gen-8B and Gen-0.6B) are currently listed as Free in the catalog; access is available anonymously or with an API key tied to a Public Cloud projectAnonymous access: 2 requests/min per IP per model. Authenticated (API key): 400 requests/min per project per model. Exceeding either returns HTTP 429card requiredOpenAI-compatible
Fireworks AI$1 trial creditVarious open-weight modelsno cardOpenAI-compatible
BasetenNew accounts receive free credits; Baseten's current pricing page does not state the amountModel APIs are priced per token; dedicated deployments are priced by compute time (per minute)OpenAI-compatible
Nebius AI Studio$1 trial credit, valid for 30 daysVarious open-weight modelscard requiredno phonecommercial OKOpenAI-compatible
Novita AI$100 Sandbox credits, valid for 90 daysVarious open-weight modelsno cardeval onlyOpenAI-compatible
AI21 Labs$10 trial credit, valid 3 monthsJamba Large/Mini: 10 RPS / 200 RPM by defaultno cardOpenAI-compatible
Alibaba Cloud (Model Studio)1,000,000 tokens (example figure, varies by model), international/Singapore region onlyQwen open & proprietary modelsno cardOpenAI-compatible
SambaNova CloudRate-limited free tier (applies when no payment method is linked) across all modelsFree Tier: 20 RPM / 20 RPD / 200,000 TPD across all models; Developer Tier (card required): 60-240 RPM depending on modelno cardcommercial OKOpenAI-compatible
Scaleway Generative APIs1,000,000 tokens free + 60 min Whisper transcription; billing starts at token 1,000,001Gemma, Llama, Mistral, Qwenno cardOpenAI-compatible
NVIDIA NIMTrial credit, phone verification requiredSome models have reduced context windows on the free trialno cardphone requiredeval onlyOpenAI-compatible
Vercel AI GatewayFree tier: $5/month credit covering a subset of models (Free Tier eligible models) at lower rate limitsFree tier is rate-limited per model (HTTP 429 on exceed), lower than paid; routes to many providers rather than hosting models itselfno cardOpenAI-compatible
Jina AI10M free tokens (one-time) across all models — embeddings, rerankers, classifier; plus a keyless Reader (r.jina.ai) for basic useFree key: 100 RPM / 100k TPM for embeddings & reranker (2 concurrent); keyless Reader 20 RPMno cardno phonecommercial OKOpenAI-compatible
AssemblyAI$50 free credit on signup (no card) — pre-recorded & streaming speech-to-text, Speech Understanding, and an OpenAI-compatible LLM Gateway (25+ models)Free tier: 5 parallel transcriptions; streaming 5 new streams/min; global 20k requests / 5 minno cardno phonecommercial OKOpenAI-compatible
ClarifaiOne-time $5 credit across serverless models (GPT-OSS-120B, Claude, Llama), vision, embeddings and image generation15 requests/second global default (CONN_THROTTLED on exceed)no cardphone requiredOpenAI-compatible
Arli AIFree plan ($0): access to all text LLMs (Gemma, Qwen, etc.), capped at ~5 requests per 2-day window, 12K context, 1 request at a time1 request at a time; ~5 requests per 2 days across all models; max 12K context; delayed responsesOpenAI-compatible
Ollama Cloud$0 Free plan: access to cloud-hosted open models (Qwen, GPT-OSS, DeepSeek, etc.) via APISession limits reset every 5 hours and weekly limits every 7 days; 1 concurrent cloud model on the free plan (exact token caps not published)no cardno phonecommercial OKOpenAI-compatible
ModelScope (API-Inference)~2,000 free API calls/day across open-weight models (Qwen3, DeepSeek, GLM, Llama, etc.) via API-Inference~2,000 calls/day; concurrency/QPS caps applied and dynamically adjustedno cardeval onlyOpenAI-compatible
Moondream Cloud$5/month usage credits in every workspace (Free plan) for the Moondream vision model — caption, query (VQA), detect, pointBounded by the $5/month creditno cardOpenAI-compatible
Sarvam AI₹100 in free credits on signup, usable across all APIs including the Sarvam-M chat/LLM API and speech (STT/TTS)Not published; bounded by the ₹100 creditno cardOpenAI-compatible
Tencent Hunyuan1,000,000 free tokens for Hunyuan text LLMs (hunyuan-a13b, turbos, translation & vision models), plus a separate 1,000,000-token allotment for hunyuan-embeddingNot published; bounded by the token packageno cardOpenAI-compatible
PoolsideLaguna M.1 (225B) and Laguna XS.2 (33B, open-weights) are free for a limited time via a self-serve developer API key (also on OpenRouter)Preview rate limits not publishedno cardOpenAI-compatible
UpstageSolar Pro 4 is free for a limited time (promo); otherwise the API is paid prepaid (commitment tiers from $100/mo) — Solar LLM (chat + embeddings) plus Document Parse / OCR / information extraction; Studio agents include 10 free runsDocument Parse billed $0.01/page (+$0.03 for extract); Tier 1 rate limits apply on prepaid tiersno cardcommercial OKOpenAI-compatible
W&B Inference$100/month of Serverless Inference credits on the Free plan (default spending cap; offer for a limited time)No RPM/TPM published; default cap of $100/month on the Free tier; concurrency limits per project/userno cardcommercial OKOpenAI-compatible
Typhoon (SCB 10X)Free to use research showcase API — all Typhoon models at $0typhoon-asr-realtime: 100 reqs/minute (documented on the ASR page); LLM/OCR models: no published quotas (beta service)no cardno phonecommercial OKOpenAI-compatible

FAQ

What does OpenAI-compatible mean?

The API accepts the same request and response format as OpenAI’s /chat/completions, so the official OpenAI SDKs work by changing only base_url and the API key.

How do I switch my code to a free OpenAI-compatible API?

Set base_url to the provider’s endpoint (listed on each provider page), use its free API key, and pass one of its free model IDs as model.

Do all free LLM APIs support the OpenAI format?

No — this list is only the providers confirmed to expose an OpenAI-compatible endpoint against their own documentation.

More guides

← All guides