Pinecone Inference

Ongoing verified 2026-08-14 · 19d ago no card

Free embeddings + rerank for RAG (5M tokens/mo)

What's free

Starter (free) plan: 5M tokens/mo for embedding models (llama-text-embed-v2, multilingual-e5-large) and 500 requests/mo for the bge-reranker-v2-m3 rerank model

Rate limits

5M embedding tokens/mo; 500 rerank requests/mo on Starter

The catch

Free rerank limited to bge-reranker-v2-m3; overages are pay-as-you-go. Not OpenAI-compatible. No card required.

TypeOngoing free tier
Free typerenewing-quota
Expiresno expiry
Modalitiesembeddings, rerank
OpenAI base URL—

Quickstart

First-party API — not OpenAI-compatible, so there is no drop-in base URL. The exact endpoint is in the official docs; authenticate with your API key in the Authorization: Bearer header.

curl -H "Authorization: Bearer $API_KEY" \
  https://<api-base-url>/<endpoint>

…or in Python:

import os, requests

r = requests.post(
    "https:///",
    headers={"Authorization": f"Bearer {os.environ['API_KEY']}"},
    json={},
)
print(r.json())

Appears in

Change history

How this free tier has changed since we started tracking it (2026-07-30) — generated from the git history of providers.json.

← All providers