An OpenAI-compatible API exposes the same /chat/completions shape, so switching is usually a one-line change: swap base_url and the key. That makes these the lowest-friction way to move side projects off paid OpenAI usage, or to add a free fallback behind the same SDK.
Each provider's exact base URL is on its page, and the free models you can pass as model are in the model index.
Top pick — Typhoon (SCB 10X)
Thai-language LLMs, OCR and speech via an OpenAI-compatible research API Details →
36 verified providers
| Provider | Free tier | Rate limits | Gotchas |
|---|---|---|---|
| Google Gemini API (AI Studio) | Gemini 2.5 Flash, 2.5 Flash-Lite, 2.5 Pro (limited), embeddings, TTS models | Per-model; Google no longer publishes the free-tier numbers in its public docs — they are shown only on the AI Studio rate-limit page after sign-in | no cardno phonecommercial OKOpenAI-compatible |
| Groq | Open-weight models (GPT-OSS, Qwen) plus Whisper, no credit card required | Free plan examples: openai/gpt-oss-120b, openai/gpt-oss-20b and qwen/qwen3.8-27b: 30 RPM/1K RPD/8K TPM/200K TPD; Whisper models: 20 RPM/2K RPD/7.2K audio seconds/hour and 28.8K/day; limits vary by model | no cardphone requiredcommercial OKOpenAI-compatible |
| OpenRouter | A rotating set of models with a :free suffix (16 on 2026-10-08; count fluctuates), single API across many providers | 20 req/min; 50 req/day under 10 credits purchased lifetime, 1000 req/day once 10+ credits purchased (one-time, not a subscription) | no cardno phonecommercial OKOpenAI-compatible |
| Cloudflare Workers AI | 10,000 Neurons/day, all account plans | 10,000 Neurons/day (free allocation). Per-task request limits: text generation 300 requests/min (models that require Workers Paid: 20 requests/min per model); text embeddings 3,000/min; speech recognition, text-to-image, translation, image-to-text 720/min | no cardno phonecommercial OKOpenAI-compatible |
| Cohere | Trial (evaluation) API keys covering chat, embed and rerank | 1,000 API calls/month total; 20 req/min chat; 2,000 inputs/min embed; 10 req/min rerank | no cardno phoneeval onlyOpenAI-compatible |
| Cerebras | $5 in free credits for new accounts, usable across all public models | Published Free Trial limits for gpt-oss-120b and qwen-3.8-27b: 5 RPM / 30,000 uncached TPM / 90,000 total TPM / 1,000,000 TPH / 1,000,000 TPD; limits vary by model | card requiredcommercial OKOpenAI-compatible |
| Mistral (La Plateforme) | "Free mode" (default for new accounts): API keys with included monthly usage and the lowest limits, "intended for evaluation and prototyping" | Not published publicly; per-model requests/second, tokens/minute and tokens/month caps are shown only in-console (Admin Panel > API > Limits) | no cardphone requiredeval onlyOpenAI-compatible |
| HuggingFace | Free CPU Basic + ZeroGPU for Spaces; Inference Providers has no included credit on the Free plan (credits must be purchased) — $2.00/mo is included only on PRO/Team/Enterprise | No RPM/TPM published, only credit amounts | no cardOpenAI-compatible |
| SiliconFlow | China platform (siliconflow.cn): several models permanently free at ¥0 (e.g. Qwen/Qwen2.5-7B-Instruct, tencent/Hunyuan-MT-7B, BAAI embeddings and rerankers). The international platform (siliconflow.com) lists no $0 LLMs and gives $1 in free credits instead | Fixed per-model limits for free models; generic docs cite ranges of 1,000-10,000 RPM and 50,000-5,000,000 TPM depending on model tier — exact limits shown in-account | no cardphone requiredOpenAI-compatible |
| Z.ai (Zhipu AI / GLM) | GLM-4.5-Flash, GLM-4.7-Flash (text), and GLM-4.6V-Flash (vision) are officially listed as $0 cost (input, cached input, and output) on a permanent basis | Not specified with concrete RPM/TPM figures in public docs | no cardno phonecommercial OKOpenAI-compatible |
| IBM watsonx.ai (Lite plan) | Lite plan: 300,000 tokens/month for foundation model inference, 20 CUH/month for ML tooling, 100 pages/month of document text extraction | 2 inference requests per second (explicitly documented for the Lite plan) | card requiredOpenAI-compatible |
| OVHcloud AI Endpoints | Seven models are currently listed as Free: two Qwen3Guard moderation models, Stable Diffusion XL and four NVIDIA Riva TTS voices; access is available anonymously or with an API key tied to a Public Cloud project | Anonymous access: 2 requests/min per IP per model. Authenticated (API key): 400 requests/min per project per model. Exceeding either returns HTTP 429 | no cardOpenAI-compatible |
| Fireworks AI | $1 trial credit | 10 requests/min with no payment method or no credits; 6,000 RPM account-wide maximum once a payment method and active credits are on file | no cardOpenAI-compatible |
| Baseten | New accounts receive free credits; Baseten's current pricing page does not state the amount | Model APIs default (models with combined token limits): Basic (unverified) 15 RPM / 100,000 TPM; Basic (verified) 120 RPM / 500,000 TPM; HTTP 429 when exceeded | OpenAI-compatible |
| Nebius AI Studio | $1 trial credit, valid for 30 days | Default per-account RPM/TPM caps are shown in the Token Factory console and in response headers (not stated numerically in the docs; the docs' 60 RPM / 400,000 TPM baseline is labeled an illustrative example). Limits auto-scale: +20% per 15-minute window at >=80% usage, up to 20x the base | card requiredno phonecommercial OKOpenAI-compatible |
| Novita AI | $100 promotional credits valid for 90 days for eligible Novita Agent Sandbox usage (CPU, RAM, storage), granted after an account-setup survey; Novita does not say they apply to Model API calls; the pricing page also lists a few LLMs as Free (e.g. Ling 3.1 Flash, Apodex 1.1 Mini) | Per model, by account tier. New/low-spend accounts are Tier T1 (monthly top-ups <= $50): typically 30 RPM and 50M TPM per model (a few models 50 RPM; some 2M-5M TPM). The Free-priced LLMs (Ling 3.1 Flash, Apodex 1.1 Mini) are 30 RPM / 50M TPM at T1 | no cardeval onlyOpenAI-compatible |
| AI21 Labs | $10 trial credit, valid 3 months | Jamba Large/Mini: 10 RPS / 200 RPM by default | no cardOpenAI-compatible |
| Alibaba Cloud (Model Studio) | Typically 1,000,000 tokens per model for new users, Singapore region (International deployment scope) only | Per-model, account-level limits that apply equally to free-quota and paid calls (International): e.g. qwen-plus 600 RPM / 1,000,000 TPM; qwen-flash and qwen-turbo 600 RPM / 5,000,000 TPM; qwen3.6-plus and qwen3.7-plus 15,000 RPM / 5,000,000 TPM. HTTP 429 when a limit is exceeded; HTTP 403 when the free quota is exhausted | no cardOpenAI-compatible |
| SambaNova Cloud | Rate-limited free tier (applies when no payment method is linked) across the models listed in the Free Tier table | Free Tier: 20 RPM / 20 RPD / 200,000 TPD listed per model in the table, enforcement scope not stated; Developer Tier (card required): 60-240 RPM depending on model | commercial OKOpenAI-compatible |
| Scaleway Generative APIs | 1,000,000 tokens free + 60 min Whisper transcription; billing starts at token 1,000,001 | With a validated payment method: 300 RPM / 200k TPM per model; 600 RPM and higher TPM once identity is verified | no cardOpenAI-compatible |
| NVIDIA NIM | Trial credit, phone verification required | Up to 40 requests/min and 10,000 requests/day; limits may vary by model and traffic from other users may cause throttling | no cardphone requiredeval onlyOpenAI-compatible |
| Vercel AI Gateway | Free tier: a monthly free credit covering a subset of models (Free Tier eligible models) at lower rate limits | No numeric free-tier limit published: the free tier has a lower per-model limit than paid (HTTP 429 when exceeded); Vercel documents behavior, not numbers | no cardOpenAI-compatible |
| Jina AI | 10M free tokens (one-time) across all models — embeddings, rerankers, classifier; plus a keyless Reader (r.jina.ai) for basic use | Free key: 100 RPM / 100k TPM for embeddings & reranker; Reader 20 RPM keyless, 500 RPM with a free key | no cardno phonecommercial OKOpenAI-compatible |
| AssemblyAI | $50 free credit on signup (no card) — pre-recorded & streaming speech-to-text, Speech Understanding, and an OpenAI-compatible LLM Gateway (25+ models) | Free tier: 5 parallel transcriptions; streaming 5 new streams/min; global 20k requests / 5 min | no cardno phonecommercial OKOpenAI-compatible |
| Clarifai | One-time $5 credit across serverless models (GPT-OSS-120B, Claude, Llama), vision, embeddings and image generation | 15 requests/second global default (CONN_THROTTLED on exceed) | no cardphone requiredOpenAI-compatible |
| Arli AI | Free plan ($0): text LLMs for testing, max 5 requests every 2 days per model, 12K context, 1 request at a time, delayed responses; no image generation | 5 requests every 2 days per model; 1 request at a time; max 12K context tokens | OpenAI-compatible |
| Ollama Cloud | $0 Free plan: starter amount of cloud usage credits, limited to a smaller set of starter models; buying credits unlocks all cloud models | 1 concurrent request on the free plan; included usage resets monthly from signup date (dollar amount of starter credits not published) | no phonecommercial OKOpenAI-compatible |
| ModelScope (API-Inference) | Free API-Inference calls paid with daily 'Magicubes': 200/day for signing in + 50/day after linking an Alibaba Cloud account (do not roll over); calls cost ~0.5 / 1 / 2 Magicubes for lightweight / standard / flagship models | Bounded by the daily Magicubes: ~250/day from signing in and linking an Alibaba Cloud account, plus any earned by other actions (~125-500 calls depending on model tier); concurrency dynamically limited, single concurrency guaranteed | no cardeval onlyOpenAI-compatible |
| Pollinations.ai | Quest Pollen: create an account, complete eligible Quests on the Quests dashboard and claim the rewards; Quest Pollen is spent before paid Pollen on regular (non-paid-only) models. No anonymous/no-key generation — every generation endpoint requires an API key and costs Pollen ($1 ≈ 1 Pollen). Text, image, video, audio, embeddings and 3D models behind one OpenAI-compatible API | No per-request rate limits published for sk_ keys; usage is bounded by the Pollen balance (legacy raw pk_ keys: 1 Pollen per IP per hour) | no cardno phoneOpenAI-compatible |
| Moondream Cloud | $5/month usage credits in every workspace (Free plan) for the Moondream vision model — caption, query (VQA), detect, point | 2 requests/sec on the free tier (10 req/sec with ≥$10 account balance); usage bounded by the $5/month credit | no cardOpenAI-compatible |
| Sarvam AI | ₹100 in free credits on signup, usable across any API including chat completion (Sarvam-105B) and speech (STT/TTS); credits never expire | Starter plan: 40 req/min for chat with Sarvam-105B (60 for other default models), STT REST and TTS REST (bulbul:v3: 30 req/min); 20 concurrent STT WebSocket streams; bounded by the ₹100 credit | commercial OKOpenAI-compatible |
| Tencent Hunyuan | One-time free package on first activation: 1,000,000 tokens shared across Hunyuan text/vision models (Hunyuan-a13b, Hunyuan-role-latest, Hunyuan-translation, Hunyuan-translation-lite, Tencent HY Vision 1.5 Instruct, Hunyuan-turbos-vision, Hunyuan-t1-vision, Hunyuan-turbos-vision-video), plus a separate 1,000,000 tokens for Hunyuan-embedding | ChatCompletions (Tencent Cloud API): default 5 concurrent requests per account; GetEmbedding: 5 requests/second (both marked pre-offline, expected 2026-12-21). A limit of 20 requests/second for the other Hunyuan APIs is not stated on any official API page checked on 2026-10-08, so it is not confirmed | no cardOpenAI-compatible |
| Poolside | Free developer API key via Poolside Platform for Poolside's Laguna coding models (current lineup: Laguna S 2.1, Laguna XS 2.1, Laguna M.1); which models and limits apply to the free key is not documented | Not published | no cardOpenAI-compatible |
| Upstage | $10 free credit on sign-up (per Upstage's Getting Started docs) for the API — Solar LLM (chat + embeddings) plus Document Parse / OCR / information extraction; Studio document agents include 10 free runs each | Tier 0 (below the Explore commitment tier): Solar Pro 4 100 RPM / 250,000 TPM; Solar Mini 4 100 RPM / 50,000 TPM; Embeddings 100 RPM / 300,000 TPM; Document OCR 3 requests/second, 500 pages/minute | no cardcommercial OKOpenAI-compatible |
| W&B Inference | Serverless Inference credits on the Free plan "for a limited time" (amount not published); Free accounts have a default spending cap of $100/month | No RPM/TPM published; per-project and per-user concurrency limits (HTTP 429 when exceeded); default spending cap $100/month on the Free tier | no cardcommercial OKOpenAI-compatible |
| Typhoon (SCB 10X) | Free to use research showcase API — all Typhoon models at $0 | typhoon-v2.5-30b-a3b-instruct: 5 requests/second, 200 requests/minute; typhoon-ocr: 2 requests/second, 20 requests/minute; typhoon-asr-realtime: 100 requests/minute | no cardno phonecommercial OKOpenAI-compatible |
FAQ
What does OpenAI-compatible mean?
The API accepts the same request and response format as OpenAI’s /chat/completions, so the official OpenAI SDKs work by changing only base_url and the API key.
How do I switch my code to a free OpenAI-compatible API?
Set base_url to the provider’s endpoint (listed on each provider page), use its free API key, and pass one of its free model IDs as model.
Do all free LLM APIs support the OpenAI format?
No — this list is only the providers confirmed to expose an OpenAI-compatible endpoint against their own documentation.