Embeddings power semantic search, RAG and clustering — and you rarely need to pay for them at prototype scale. The providers below expose embedding models on a free tier, several with generous monthly token allowances.
Building RAG? Pair one of these with a free chat model (see the model index), ideally behind the same OpenAI-compatible base URL.
Top pick — Jina AI
Free embeddings & rerank behind an OpenAI-compatible endpoint Details →
13 verified providers
| Provider | Free tier | Rate limits | Gotchas |
|---|---|---|---|
| Google Gemini API (AI Studio) | Gemini 2.5 Flash, 2.5 Flash-Lite, 2.5 Pro (limited), embeddings, TTS models | Per-model; Google no longer publishes the free-tier numbers in its public docs — they are shown only on the AI Studio rate-limit page after sign-in | no cardno phonecommercial OKOpenAI-compatible |
| Cloudflare Workers AI | 10,000 Neurons/day, all account plans | 10,000 Neurons/day (free allocation). Per-task request limits: text generation 300 requests/min (models that require Workers Paid: 20 requests/min per model); text embeddings 3,000/min; speech recognition, text-to-image, translation, image-to-text 720/min | no cardno phonecommercial OKOpenAI-compatible |
| Cohere | Trial (evaluation) API keys covering chat, embed and rerank | 1,000 API calls/month total; 20 req/min chat; 2,000 inputs/min embed; 10 req/min rerank | no cardno phoneeval onlyOpenAI-compatible |
| HuggingFace | Free CPU Basic + ZeroGPU for Spaces; Inference Providers has no included credit on the Free plan (credits must be purchased) — $2.00/mo is included only on PRO/Team/Enterprise | No RPM/TPM published, only credit amounts | no cardOpenAI-compatible |
| SiliconFlow | China platform (siliconflow.cn): several models permanently free at ¥0 (e.g. Qwen/Qwen2.5-7B-Instruct, tencent/Hunyuan-MT-7B, BAAI embeddings and rerankers). The international platform (siliconflow.com) lists no $0 LLMs and gives $1 in free credits instead | Fixed per-model limits for free models; generic docs cite ranges of 1,000-10,000 RPM and 50,000-5,000,000 TPM depending on model tier — exact limits shown in-account | no cardphone requiredOpenAI-compatible |
| Jina AI | 10M free tokens (one-time) across all models — embeddings, rerankers, classifier; plus a keyless Reader (r.jina.ai) for basic use | Free key: 100 RPM / 100k TPM for embeddings & reranker; Reader 20 RPM keyless, 500 RPM with a free key | no cardno phonecommercial OKOpenAI-compatible |
| Mixedbread | Starter plan: $5 one-time credit (no card) — embeddings (mxbai-embed-large-v1), reranking, and multimodal search over PDF/image/doc/code | Starter plan: 100 requests/min | no cardno phonecommercial OK |
| Clarifai | One-time $5 credit across serverless models (GPT-OSS-120B, Claude, Llama), vision, embeddings and image generation | 15 requests/second global default (CONN_THROTTLED on exceed) | no cardphone requiredOpenAI-compatible |
| Pinecone Inference | Starter (free) plan: 5M embedding tokens/month per model (llama-text-embed-v2, multilingual-e5-large, pinecone-sparse-english-v0) and 500 rerank requests/month (bge-reranker-v2-m3; the rate-limits doc also lists pinecone-rerank-v0 at 500) | Starter: 5M embedding tokens/mo per model; 250K embedding tokens/min per model (passage); 500 rerank requests/mo and 60 rerank requests/min per model; 100 inference requests/s and 2,000/min per project | no card |
| Twelve Labs (Marengo Embed) | Free plan: the Marengo 3.5 Embed API (video, audio, image, text and document inputs) and the Pegasus Analyze API at no cost, plus 600 minutes (10 hours) of video indexing in total | Embed video/audio: 3,000 RPD, 25 RPM; embed text/image: 3,000 RPD, 600 RPM | no card |
| Tencent Hunyuan | One-time free package on first activation: 1,000,000 tokens shared across Hunyuan text/vision models (Hunyuan-a13b, Hunyuan-role-latest, Hunyuan-translation, Hunyuan-translation-lite, Tencent HY Vision 1.5 Instruct, Hunyuan-turbos-vision, Hunyuan-t1-vision, Hunyuan-turbos-vision-video), plus a separate 1,000,000 tokens for Hunyuan-embedding | ChatCompletions (Tencent Cloud API): default 5 concurrent requests per account; GetEmbedding: 5 requests/second (both marked pre-offline, expected 2026-12-21). A limit of 20 requests/second for the other Hunyuan APIs is not stated on any official API page checked on 2026-10-08, so it is not confirmed | no cardOpenAI-compatible |
| Voyage AI | 200M free tokens per model on current embedding models (voyage-4-large, voyage-4, voyage-4-lite, voyage-context-4, voyage-code-4) and on the rerankers listed in the price table (rerank-3, rerank-3-lite); voyage-multimodal-3.5 and voyage-multimodal-3 get 200M text tokens + 150B pixels; legacy voyage-finance-2, voyage-law-2 and voyage-code-2 get 50M — a one-time complimentary allotment per account | No numeric limit published for accounts without a payment method. With a payment method (Tier 1, free tokens still apply), per model: voyage-4 / voyage-code-4 8M TPM, 2,000 RPM; voyage-4-large / voyage-context-4 3M TPM, 2,000 RPM; voyage-4-lite 16M TPM, 2,000 RPM; voyage-multimodal-3.5 2M TPM, 2,000 RPM; rerank-3 / rerank-2.5 2M TPM, 2,000 RPM; rerank-3-lite / rerank-2.5-lite 4M TPM, 2,000 RPM. HTTP 429 on exceed | no card |
| Upstage | $10 free credit on sign-up (per Upstage's Getting Started docs) for the API — Solar LLM (chat + embeddings) plus Document Parse / OCR / information extraction; Studio document agents include 10 free runs each | Tier 0 (below the Explore commitment tier): Solar Pro 4 100 RPM / 250,000 TPM; Solar Mini 4 100 RPM / 50,000 TPM; Embeddings 100 RPM / 300,000 TPM; Document OCR 3 requests/second, 500 pages/minute | no cardcommercial OKOpenAI-compatible |
FAQ
Is there a free embeddings API?
Yes — providers such as Jina AI, Cohere, Google Gemini and Cloudflare offer embedding models on a free tier, several with large monthly token allowances.
Can I build RAG for free?
For prototypes, yes: pair a free embeddings API with a free chat model, ideally behind the same OpenAI-compatible base URL.