The best open-weight models. One API.
Connect Claude Code, Cline, opencode or your application to Kimi, DeepSeek, Qwen and more through an OpenAI and Anthropic compatible API.
- Chat & code
- Audio → text
- Image generation
No account
# just point your base_url
base_url = "https://api.velqa.dev/v1"
api_key = "sk-..."
model = "kimi-k2.6"Your entire AI stack, without multiplying integrations.
Code, audio, embeddings and images run through the same infrastructure. One key, one view and one consistent experience.

Available models
Open-weight only — no closed source.
Raisonnement lourd et agentique, contexte 1M — lent, hors forfait
Flagship 2.4T — code et travail pro, hors forfait
Flagship raisonnement lourd
Embeddings multilingues (FR / AR) — RAG et recherche sémantique
Transcription audio multilingue (arabe / darija / FR) — facturé à la minute
Génération d'images texte → image, multilingue — facturé à l'image
How a request flows
One gateway, transparent routing to the best available model.
Your request
OpenAI or Anthropic SDK, base_url pointed at Velqa.
Velqa gateway
Checks your quota, picks the available model.
Automatic switch
Quota exceeded → pay-as-you-go, no interruption.
Response
Same format as direct, stable latency from the EU.
Agents that execute, not just models that answer.
The Sandbox runs Python or Node.js code in an isolated environment. Your agent reads files, executes commands and iterates on its own — every turn metered and billed for real. Included on all three plans: 1h to 15h of compute per month, then $0.30/h.
Python 3 & Node.js runtimes
Pick the environment, add your instructions, and the agent boots in seconds.
Per-agent isolation
Every agent runs in its own sandbox, walled off from the rest. Nothing leaks between runs.
Billed per turn
Compute is metered and charged turn by turn, visible in your dashboard. No idle cost.
analyze-csv
Python 3 · isolated
Why Velqa
Built for demanding developers, with no technical compromise.
Pay by card
Stripe. No hidden fees.
OpenAI & Anthropic compatible
One base_url, two API formats. Your existing SDK and tools work without changes.
Transparent quotas
Fixed-reset sessions (5h and 7 days), real-time consumption in the dashboard. You always know where you stand.
No hard cutoff
Quota exceeded? Automatic switch to pay-as-you-go. Your agent keeps running, you stay in control.
Hosted in Europe
Gateway in the EU, stable and predictable latency. Your prompts don't pass through opaque relays.
Encrypted keys
AES-256-GCM at rest, one-click rotation, budgets and limits per key. Security isn't optional.
PAGE AGENT
An assistant that understands your page.
Page Agent reads your site's visible context to answer, guide visitors, and prepare useful actions.
- Understands your page context
- Confirmation before sensitive actions
- Masked and protected data
- Bounded and controlled execution
A cockpit, not a black box.
Track requests, session quota and budget from one screen. Important signals stay visible before they become problems.
Explore the dashboardUp and running in 3 steps
From account to first request in under 2 minutes.
- 01
Create your account
Sign up in 30 seconds. Choose a plan and pay via Stripe.
- 02
Generate your API key
From the dashboard, an sk-... key with dedicated budget and limits. Rotate anytime.
- 03
Point your base_url
Claude Code, opencode, Cline, Roo Code or your own code: change one config line and go.
Pricing
Prepaid subscription. Annual: −8% on Dev and Pro.
Prices in USD, taxes not included. Payment via Stripe.
Cancel anytime.
Starter
For testing an idea or a side project
- Up to 70M tokens / month
- ~307 prompts / 5 h, ~615 / week
- DeepSeek V4 Flash, GLM-4.7 Flash, GLM-4.7 and Hunyuan Hy3
- Sandbox agent — ~1h included / month
- Media library: 7 days
- Community support
Dev
For coding every day with your agent
- Up to 140M tokens / month
- ~789 prompts / 5 h, ~1579 / week
- Everything in Starter, + MiniMax M3, MiMo V2.5, Kimi K2.6 and GLM-5.2
- Sandbox agent — ~5h included / month
- Media library: 30 days
- Email support
Pro
For heavy individual use and premium models
- Up to 290M tokens / month
- ~878 prompts / 5 h, ~1757 / week
- Everything in Dev, + DeepSeek V4 Pro and Qwen3.7 Max
- Sandbox agent — ~15h included / month
- Media library: 90 days
- Priority support
Quota is a $ budget, not a fixed token count: it lasts longer on an economical model. Exceeding it → automatic Recharge Boost switch. No hard cutoff.
Agent Sandbox included on all three plans: 1h to 15h of compute per month depending on the plan. Beyond that, $0.30/h drawn from your Boost balance — or an hour pack at a lower rate.
Need a custom usage plan? Contact us.
Works with your tools
Just point your base_url. Nothing else.
# Cline — settings.json
"openAIBaseURL": "https://api.velqa.dev/v1",
"openAIApiKey": "sk-..."Frequently asked questions
Another question? Email us, we reply fast.
What payment methods do you accept?
International cards via Stripe. Prices are shown in USD, taxes not included.
Is the API really compatible with OpenAI and Anthropic?
Yes. The /v1 endpoint accepts the OpenAI format (chat/completions) and the Anthropic format (messages). Any SDK or tool that lets you change the base_url works out of the box.
What happens if I exceed my quota?
No interruption: requests automatically switch to pay-as-you-go, debited from your balance. You can also disable this for your account.
Which models do you offer?
Only state-of-the-art open-weight models: Kimi K3, Qwen3.8 Max, DeepSeek V4, GLM-4.7, MiniMax M3… plus embedding models (BGE-M3, Qwen3 Embedding), audio transcription (Whisper Large v3), and image generation (FLUX.2). The list evolves monthly based on benchmarks.
Can I cancel my subscription?
Yes, anytime from the dashboard. Your plan stays active until the end of the paid period and does not renew.
Where is your infrastructure hosted?
The gateway and database are hosted in the European Union. API keys are encrypted AES-256-GCM at rest.
Ready to code with the best models?
Free account, first API key in a few clicks.


