Free embeddings APIs for search and RAG

Free vector embeddings for semantic search, RAG and clustering.

Embeddings power semantic search, RAG and clustering — and you rarely need to pay for them at prototype scale. The providers below expose embedding models on a free tier, several with generous monthly token allowances.

Building RAG? Pair one of these with a free chat model (see the model index), ideally behind the same OpenAI-compatible base URL.

Top pick — Jina AI

Free embeddings & rerank behind an OpenAI-compatible endpoint Details →

13 verified providers

ProviderFree tierRate limitsGotchas
Google Gemini API (AI Studio)Gemini 2.5 Flash, 2.5 Flash-Lite, 2.5 Pro (limited), embeddings, TTS modelsPer-model; Google no longer publishes the free-tier numbers in its public docs — they are shown only on the AI Studio rate-limit page after sign-inno cardno phonecommercial OKOpenAI-compatible
Cloudflare Workers AI10,000 Neurons/day, all account plans10,000 Neurons/day (free allocation). Per-task request limits: text generation 300 requests/min (models that require Workers Paid: 20 requests/min per model); text embeddings 3,000/min; speech recognition, text-to-image, translation, image-to-text 720/minno cardno phonecommercial OKOpenAI-compatible
CohereTrial (evaluation) API keys covering chat, embed and rerank1,000 API calls/month total; 20 req/min chat; 2,000 inputs/min embed; 10 req/min rerankno cardno phoneeval onlyOpenAI-compatible
HuggingFaceFree CPU Basic + ZeroGPU for Spaces; Inference Providers has no included credit on the Free plan (credits must be purchased) — $2.00/mo is included only on PRO/Team/EnterpriseNo RPM/TPM published, only credit amountsno cardOpenAI-compatible
SiliconFlowChina platform (siliconflow.cn): several models permanently free at ¥0 (e.g. Qwen/Qwen2.5-7B-Instruct, tencent/Hunyuan-MT-7B, BAAI embeddings and rerankers). The international platform (siliconflow.com) lists no $0 LLMs and gives $1 in free credits insteadFixed per-model limits for free models; generic docs cite ranges of 1,000-10,000 RPM and 50,000-5,000,000 TPM depending on model tier — exact limits shown in-accountno cardphone requiredOpenAI-compatible
Jina AI10M free tokens (one-time) across all models — embeddings, rerankers, classifier; plus a keyless Reader (r.jina.ai) for basic useFree key: 100 RPM / 100k TPM for embeddings & reranker; Reader 20 RPM keyless, 500 RPM with a free keyno cardno phonecommercial OKOpenAI-compatible
MixedbreadStarter plan: $5 one-time credit (no card) — embeddings (mxbai-embed-large-v1), reranking, and multimodal search over PDF/image/doc/codeStarter plan: 100 requests/minno cardno phonecommercial OK
ClarifaiOne-time $5 credit across serverless models (GPT-OSS-120B, Claude, Llama), vision, embeddings and image generation15 requests/second global default (CONN_THROTTLED on exceed)no cardphone requiredOpenAI-compatible
Pinecone InferenceStarter (free) plan: 5M embedding tokens/month per model (llama-text-embed-v2, multilingual-e5-large, pinecone-sparse-english-v0) and 500 rerank requests/month (bge-reranker-v2-m3; the rate-limits doc also lists pinecone-rerank-v0 at 500)Starter: 5M embedding tokens/mo per model; 250K embedding tokens/min per model (passage); 500 rerank requests/mo and 60 rerank requests/min per model; 100 inference requests/s and 2,000/min per projectno card
Twelve Labs (Marengo Embed)Free plan: the Marengo 3.5 Embed API (video, audio, image, text and document inputs) and the Pegasus Analyze API at no cost, plus 600 minutes (10 hours) of video indexing in totalEmbed video/audio: 3,000 RPD, 25 RPM; embed text/image: 3,000 RPD, 600 RPMno card
Tencent HunyuanOne-time free package on first activation: 1,000,000 tokens shared across Hunyuan text/vision models (Hunyuan-a13b, Hunyuan-role-latest, Hunyuan-translation, Hunyuan-translation-lite, Tencent HY Vision 1.5 Instruct, Hunyuan-turbos-vision, Hunyuan-t1-vision, Hunyuan-turbos-vision-video), plus a separate 1,000,000 tokens for Hunyuan-embeddingChatCompletions (Tencent Cloud API): default 5 concurrent requests per account; GetEmbedding: 5 requests/second (both marked pre-offline, expected 2026-12-21). A limit of 20 requests/second for the other Hunyuan APIs is not stated on any official API page checked on 2026-10-08, so it is not confirmedno cardOpenAI-compatible
Voyage AI200M free tokens per model on current embedding models (voyage-4-large, voyage-4, voyage-4-lite, voyage-context-4, voyage-code-4) and on the rerankers listed in the price table (rerank-3, rerank-3-lite); voyage-multimodal-3.5 and voyage-multimodal-3 get 200M text tokens + 150B pixels; legacy voyage-finance-2, voyage-law-2 and voyage-code-2 get 50M — a one-time complimentary allotment per accountNo numeric limit published for accounts without a payment method. With a payment method (Tier 1, free tokens still apply), per model: voyage-4 / voyage-code-4 8M TPM, 2,000 RPM; voyage-4-large / voyage-context-4 3M TPM, 2,000 RPM; voyage-4-lite 16M TPM, 2,000 RPM; voyage-multimodal-3.5 2M TPM, 2,000 RPM; rerank-3 / rerank-2.5 2M TPM, 2,000 RPM; rerank-3-lite / rerank-2.5-lite 4M TPM, 2,000 RPM. HTTP 429 on exceedno card
Upstage$10 free credit on sign-up (per Upstage's Getting Started docs) for the API — Solar LLM (chat + embeddings) plus Document Parse / OCR / information extraction; Studio document agents include 10 free runs eachTier 0 (below the Explore commitment tier): Solar Pro 4 100 RPM / 250,000 TPM; Solar Mini 4 100 RPM / 50,000 TPM; Embeddings 100 RPM / 300,000 TPM; Document OCR 3 requests/second, 500 pages/minuteno cardcommercial OKOpenAI-compatible

FAQ

Is there a free embeddings API?

Yes — providers such as Jina AI, Cohere, Google Gemini and Cloudflare offer embedding models on a free tier, several with large monthly token allowances.

Can I build RAG for free?

For prototypes, yes: pair a free embeddings API with a free chat model, ideally behind the same OpenAI-compatible base URL.

More guides

← All guides