Free OpenAI-compatible APIs (drop-in base_url)

Point the OpenAI SDK at a different base_url and keep all of your existing code.

An OpenAI-compatible API exposes the same /chat/completions shape, so switching is usually a one-line change: swap base_url and the key. That makes these the lowest-friction way to move side projects off paid OpenAI usage, or to add a free fallback behind the same SDK.

Each provider's exact base URL is on its page, and the free models you can pass as model are in the model index.

Top pick — Typhoon (SCB 10X)

Thai-language LLMs, OCR and speech via an OpenAI-compatible research API Details →

36 verified providers

ProviderFree tierRate limitsGotchas
Google Gemini API (AI Studio)Gemini 2.5 Flash, 2.5 Flash-Lite, 2.5 Pro (limited), embeddings, TTS modelsPer-model; Google no longer publishes the free-tier numbers in its public docs — they are shown only on the AI Studio rate-limit page after sign-inno cardno phonecommercial OKOpenAI-compatible
GroqOpen-weight models (GPT-OSS, Qwen) plus Whisper, no credit card requiredFree plan examples: openai/gpt-oss-120b, openai/gpt-oss-20b and qwen/qwen3.8-27b: 30 RPM/1K RPD/8K TPM/200K TPD; Whisper models: 20 RPM/2K RPD/7.2K audio seconds/hour and 28.8K/day; limits vary by modelno cardphone requiredcommercial OKOpenAI-compatible
OpenRouterA rotating set of models with a :free suffix (16 on 2026-10-08; count fluctuates), single API across many providers20 req/min; 50 req/day under 10 credits purchased lifetime, 1000 req/day once 10+ credits purchased (one-time, not a subscription)no cardno phonecommercial OKOpenAI-compatible
Cloudflare Workers AI10,000 Neurons/day, all account plans10,000 Neurons/day (free allocation). Per-task request limits: text generation 300 requests/min (models that require Workers Paid: 20 requests/min per model); text embeddings 3,000/min; speech recognition, text-to-image, translation, image-to-text 720/minno cardno phonecommercial OKOpenAI-compatible
CohereTrial (evaluation) API keys covering chat, embed and rerank1,000 API calls/month total; 20 req/min chat; 2,000 inputs/min embed; 10 req/min rerankno cardno phoneeval onlyOpenAI-compatible
Cerebras$5 in free credits for new accounts, usable across all public modelsPublished Free Trial limits for gpt-oss-120b and qwen-3.8-27b: 5 RPM / 30,000 uncached TPM / 90,000 total TPM / 1,000,000 TPH / 1,000,000 TPD; limits vary by modelcard requiredcommercial OKOpenAI-compatible
Mistral (La Plateforme)"Free mode" (default for new accounts): API keys with included monthly usage and the lowest limits, "intended for evaluation and prototyping"Not published publicly; per-model requests/second, tokens/minute and tokens/month caps are shown only in-console (Admin Panel > API > Limits)no cardphone requiredeval onlyOpenAI-compatible
HuggingFaceFree CPU Basic + ZeroGPU for Spaces; Inference Providers has no included credit on the Free plan (credits must be purchased) — $2.00/mo is included only on PRO/Team/EnterpriseNo RPM/TPM published, only credit amountsno cardOpenAI-compatible
SiliconFlowChina platform (siliconflow.cn): several models permanently free at ¥0 (e.g. Qwen/Qwen2.5-7B-Instruct, tencent/Hunyuan-MT-7B, BAAI embeddings and rerankers). The international platform (siliconflow.com) lists no $0 LLMs and gives $1 in free credits insteadFixed per-model limits for free models; generic docs cite ranges of 1,000-10,000 RPM and 50,000-5,000,000 TPM depending on model tier — exact limits shown in-accountno cardphone requiredOpenAI-compatible
Z.ai (Zhipu AI / GLM)GLM-4.5-Flash, GLM-4.7-Flash (text), and GLM-4.6V-Flash (vision) are officially listed as $0 cost (input, cached input, and output) on a permanent basisNot specified with concrete RPM/TPM figures in public docsno cardno phonecommercial OKOpenAI-compatible
IBM watsonx.ai (Lite plan)Lite plan: 300,000 tokens/month for foundation model inference, 20 CUH/month for ML tooling, 100 pages/month of document text extraction2 inference requests per second (explicitly documented for the Lite plan)card requiredOpenAI-compatible
OVHcloud AI EndpointsSeven models are currently listed as Free: two Qwen3Guard moderation models, Stable Diffusion XL and four NVIDIA Riva TTS voices; access is available anonymously or with an API key tied to a Public Cloud projectAnonymous access: 2 requests/min per IP per model. Authenticated (API key): 400 requests/min per project per model. Exceeding either returns HTTP 429no cardOpenAI-compatible
Fireworks AI$1 trial credit10 requests/min with no payment method or no credits; 6,000 RPM account-wide maximum once a payment method and active credits are on fileno cardOpenAI-compatible
BasetenNew accounts receive free credits; Baseten's current pricing page does not state the amountModel APIs default (models with combined token limits): Basic (unverified) 15 RPM / 100,000 TPM; Basic (verified) 120 RPM / 500,000 TPM; HTTP 429 when exceededOpenAI-compatible
Nebius AI Studio$1 trial credit, valid for 30 daysDefault per-account RPM/TPM caps are shown in the Token Factory console and in response headers (not stated numerically in the docs; the docs' 60 RPM / 400,000 TPM baseline is labeled an illustrative example). Limits auto-scale: +20% per 15-minute window at >=80% usage, up to 20x the basecard requiredno phonecommercial OKOpenAI-compatible
Novita AI$100 promotional credits valid for 90 days for eligible Novita Agent Sandbox usage (CPU, RAM, storage), granted after an account-setup survey; Novita does not say they apply to Model API calls; the pricing page also lists a few LLMs as Free (e.g. Ling 3.1 Flash, Apodex 1.1 Mini)Per model, by account tier. New/low-spend accounts are Tier T1 (monthly top-ups <= $50): typically 30 RPM and 50M TPM per model (a few models 50 RPM; some 2M-5M TPM). The Free-priced LLMs (Ling 3.1 Flash, Apodex 1.1 Mini) are 30 RPM / 50M TPM at T1no cardeval onlyOpenAI-compatible
AI21 Labs$10 trial credit, valid 3 monthsJamba Large/Mini: 10 RPS / 200 RPM by defaultno cardOpenAI-compatible
Alibaba Cloud (Model Studio)Typically 1,000,000 tokens per model for new users, Singapore region (International deployment scope) onlyPer-model, account-level limits that apply equally to free-quota and paid calls (International): e.g. qwen-plus 600 RPM / 1,000,000 TPM; qwen-flash and qwen-turbo 600 RPM / 5,000,000 TPM; qwen3.6-plus and qwen3.7-plus 15,000 RPM / 5,000,000 TPM. HTTP 429 when a limit is exceeded; HTTP 403 when the free quota is exhaustedno cardOpenAI-compatible
SambaNova CloudRate-limited free tier (applies when no payment method is linked) across the models listed in the Free Tier tableFree Tier: 20 RPM / 20 RPD / 200,000 TPD listed per model in the table, enforcement scope not stated; Developer Tier (card required): 60-240 RPM depending on modelcommercial OKOpenAI-compatible
Scaleway Generative APIs1,000,000 tokens free + 60 min Whisper transcription; billing starts at token 1,000,001With a validated payment method: 300 RPM / 200k TPM per model; 600 RPM and higher TPM once identity is verifiedno cardOpenAI-compatible
NVIDIA NIMTrial credit, phone verification requiredUp to 40 requests/min and 10,000 requests/day; limits may vary by model and traffic from other users may cause throttlingno cardphone requiredeval onlyOpenAI-compatible
Vercel AI GatewayFree tier: a monthly free credit covering a subset of models (Free Tier eligible models) at lower rate limitsNo numeric free-tier limit published: the free tier has a lower per-model limit than paid (HTTP 429 when exceeded); Vercel documents behavior, not numbersno cardOpenAI-compatible
Jina AI10M free tokens (one-time) across all models — embeddings, rerankers, classifier; plus a keyless Reader (r.jina.ai) for basic useFree key: 100 RPM / 100k TPM for embeddings & reranker; Reader 20 RPM keyless, 500 RPM with a free keyno cardno phonecommercial OKOpenAI-compatible
AssemblyAI$50 free credit on signup (no card) — pre-recorded & streaming speech-to-text, Speech Understanding, and an OpenAI-compatible LLM Gateway (25+ models)Free tier: 5 parallel transcriptions; streaming 5 new streams/min; global 20k requests / 5 minno cardno phonecommercial OKOpenAI-compatible
ClarifaiOne-time $5 credit across serverless models (GPT-OSS-120B, Claude, Llama), vision, embeddings and image generation15 requests/second global default (CONN_THROTTLED on exceed)no cardphone requiredOpenAI-compatible
Arli AIFree plan ($0): text LLMs for testing, max 5 requests every 2 days per model, 12K context, 1 request at a time, delayed responses; no image generation5 requests every 2 days per model; 1 request at a time; max 12K context tokensOpenAI-compatible
Ollama Cloud$0 Free plan: starter amount of cloud usage credits, limited to a smaller set of starter models; buying credits unlocks all cloud models1 concurrent request on the free plan; included usage resets monthly from signup date (dollar amount of starter credits not published)no phonecommercial OKOpenAI-compatible
ModelScope (API-Inference)Free API-Inference calls paid with daily 'Magicubes': 200/day for signing in + 50/day after linking an Alibaba Cloud account (do not roll over); calls cost ~0.5 / 1 / 2 Magicubes for lightweight / standard / flagship modelsBounded by the daily Magicubes: ~250/day from signing in and linking an Alibaba Cloud account, plus any earned by other actions (~125-500 calls depending on model tier); concurrency dynamically limited, single concurrency guaranteedno cardeval onlyOpenAI-compatible
Pollinations.aiQuest Pollen: create an account, complete eligible Quests on the Quests dashboard and claim the rewards; Quest Pollen is spent before paid Pollen on regular (non-paid-only) models. No anonymous/no-key generation — every generation endpoint requires an API key and costs Pollen ($1 ≈ 1 Pollen). Text, image, video, audio, embeddings and 3D models behind one OpenAI-compatible APINo per-request rate limits published for sk_ keys; usage is bounded by the Pollen balance (legacy raw pk_ keys: 1 Pollen per IP per hour)no cardno phoneOpenAI-compatible
Moondream Cloud$5/month usage credits in every workspace (Free plan) for the Moondream vision model — caption, query (VQA), detect, point2 requests/sec on the free tier (10 req/sec with ≥$10 account balance); usage bounded by the $5/month creditno cardOpenAI-compatible
Sarvam AI₹100 in free credits on signup, usable across any API including chat completion (Sarvam-105B) and speech (STT/TTS); credits never expireStarter plan: 40 req/min for chat with Sarvam-105B (60 for other default models), STT REST and TTS REST (bulbul:v3: 30 req/min); 20 concurrent STT WebSocket streams; bounded by the ₹100 creditcommercial OKOpenAI-compatible
Tencent HunyuanOne-time free package on first activation: 1,000,000 tokens shared across Hunyuan text/vision models (Hunyuan-a13b, Hunyuan-role-latest, Hunyuan-translation, Hunyuan-translation-lite, Tencent HY Vision 1.5 Instruct, Hunyuan-turbos-vision, Hunyuan-t1-vision, Hunyuan-turbos-vision-video), plus a separate 1,000,000 tokens for Hunyuan-embeddingChatCompletions (Tencent Cloud API): default 5 concurrent requests per account; GetEmbedding: 5 requests/second (both marked pre-offline, expected 2026-12-21). A limit of 20 requests/second for the other Hunyuan APIs is not stated on any official API page checked on 2026-10-08, so it is not confirmedno cardOpenAI-compatible
PoolsideFree developer API key via Poolside Platform for Poolside's Laguna coding models (current lineup: Laguna S 2.1, Laguna XS 2.1, Laguna M.1); which models and limits apply to the free key is not documentedNot publishedno cardOpenAI-compatible
Upstage$10 free credit on sign-up (per Upstage's Getting Started docs) for the API — Solar LLM (chat + embeddings) plus Document Parse / OCR / information extraction; Studio document agents include 10 free runs eachTier 0 (below the Explore commitment tier): Solar Pro 4 100 RPM / 250,000 TPM; Solar Mini 4 100 RPM / 50,000 TPM; Embeddings 100 RPM / 300,000 TPM; Document OCR 3 requests/second, 500 pages/minuteno cardcommercial OKOpenAI-compatible
W&B InferenceServerless Inference credits on the Free plan "for a limited time" (amount not published); Free accounts have a default spending cap of $100/monthNo RPM/TPM published; per-project and per-user concurrency limits (HTTP 429 when exceeded); default spending cap $100/month on the Free tierno cardcommercial OKOpenAI-compatible
Typhoon (SCB 10X)Free to use research showcase API — all Typhoon models at $0typhoon-v2.5-30b-a3b-instruct: 5 requests/second, 200 requests/minute; typhoon-ocr: 2 requests/second, 20 requests/minute; typhoon-asr-realtime: 100 requests/minuteno cardno phonecommercial OKOpenAI-compatible

FAQ

What does OpenAI-compatible mean?

The API accepts the same request and response format as OpenAI’s /chat/completions, so the official OpenAI SDKs work by changing only base_url and the API key.

How do I switch my code to a free OpenAI-compatible API?

Set base_url to the provider’s endpoint (listed on each provider page), use its free API key, and pass one of its free model IDs as model.

Do all free LLM APIs support the OpenAI format?

No — this list is only the providers confirmed to expose an OpenAI-compatible endpoint against their own documentation.

More guides

← All guides