Reranking with Velqa.dev — reorder your RAG results (FR / AR)
Velqa offers multilingual reranking via qwen3-reranker-8b (Alibaba), on the API https://api.velqa.dev/v1/rerank. The reranker takes a query and a list of documents, then reorders them by their real relevance to the query — it complements embeddings in a RAG pipeline to surface the best passages before sending them to the chat model. It is included in every plan (Starter, Dev, Pro, Recharge Boost).
Why a reranker?
Embedding search is fast but approximate: it brings vectors close together, not necessarily exact meaning. A reranker re-reads each (query, document) pair and assigns a fine-grained relevance_score. The typical pattern:
- Embeddings (
bge-m3orqwen3-embedding-8b) fetch the 20–50 closest candidates. qwen3-reranker-8breorders them and you keep the top 3–5.- You send that top set to the chat model as context.
Pricing
Reranking is billed on input tokens (query + documents), with no output tokens. Current USD pricing and checkout availability are shown in the dashboard; checkout is Stripe, in USD, by international card.
Example with curl
curl https://api.velqa.dev/v1/rerank \
-H "Authorization: Bearer $VELQA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-reranker-8b",
"query": "How do I add prepaid credit to my account?",
"documents": [
"Use Stripe Checkout to add USD prepaid credit.",
"Whisper transcribes audio in French and darija.",
"Choose Top up, then select a USD amount."
]
}'Example with Python
import requests
resp = requests.post(
"https://api.velqa.dev/v1/rerank",
headers={"Authorization": "Bearer VELQA_API_KEY"},
json={
"model": "qwen3-reranker-8b",
"query": "How do I add prepaid credit to my account?",
"documents": [
"Use Stripe Checkout to add USD prepaid credit.",
"Whisper transcribes audio in French and darija.",
"Choose Top up, then select a USD amount.",
],
},
)
for r in resp.json()["results"]:
print(round(r["relevance_score"], 3), r["index"])The response contains results, a list sorted from most to least relevant, where each item exposes index (the document's position in your request) and relevance_score. Reorder your documents using index and keep the best ones for your context.
See also
- RAG pipeline — the full end-to-end pipeline, with pgvector.
- Embeddings — the first step of the RAG pipeline.
- Available models — full catalog and current availability.
